Preprint
Sep 2026
Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement
Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency.
Jia-Rui Yang, Jia-Jin Zhang, Bin Zhu et al.
· 0 citations