Vision-language-action (VLA) models handle long-horizon manipulation, yet success hinges on a few precision-critical phases where millimeter-scale errors undo all prior progress. Online reinforcement learning (RL) can optimize exactly these actions, but free exploration is far too costly on real robots, which makes hum...
Wei-Hui Zhao, Xiao Yan, Zu-Nian Wan et al.· 0 citations
This work introduces Continual Interactive Distillation for Embodied Reinforcement Learning (CIDER), a continual reinforcement learning framework that freezes the accumulated historical policy as a teacher before learning each new task and interleaves task learning with distillation-based retention.
Hou-Lin Li, Ming Xu, Guofeng Xu et al.· 0 citations
VINE is proposed, an RL-oriented sampling method that enables stable end-to-end value-gradient optimization for flow-matching policies and achieves stable policy improvement and consistently outperforms state-of-the-art RL methods on the OGBench offline RL benchmark and real-world robotic manipulation task.
Rushuai Yang, Zhuo Han, Houlin Li et al.· arXiv.org· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.