Review
Memory-Augmented VLM Planners for Long-Horizon VLA Control via RL
This work proposes a demo-free hierarchical memory VLA : a Qwen3-VL-4B planner with a persistent keyframe buffer, trained via streaming GRPO on dense task-completion reward from simulation, above a frozen GroundSG π 0 .
Krish Sharma, Lucas Burgett, Sharma Burgett
· 0 citations