Skip to content
Preprint

Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement

Sep 2026 · 0 citations · 31 references
Computer Science

TL;DR

Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency.

Abstract

Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flow actor reuses the initial noise stored during rollout and learns through advantage-weighted conditional likelihood, directly supervising the action mapping used for execution. We evaluate RAPolicy across four single-task settings and one joint five-task setting in the real world, with online training budgets of only 1--2 hours. Starting from policies fine-tuned on just 10 demonstrations per task, RAPolicy rapidly adapts to new single tasks and achieves an average 86.3% success rate. In the joint multi-task setting, RAPolicy improves overall success rate from 52% to 88% while preserving performance on already reliable tasks and improving weaker capabilities. Overall, RAPolicy substantially outperforms the baselines in aggregate task success while requiring fewer human interventions, demonstrating stable policy improvement and high online training efficiency. Project page: https://flyfaerss.github.io/RAPolicy.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

VLA-Precision: Asymmetric Co-Bootstrapping for Efficient Real-World Online RL of Vision-Language-Action Models

VLA-Precision, an efficient real-world online RL framework featuring the Asymmetric Co-Bootstrapping (ACoB) algorithm and the ACoB-Stream architecture, and ACoB-Stream, a closed-loop experience--policy architecture that establishes invariant-state decoupling and on-demand streaming as design principles, deliver up to 1...

Chen-Yu Su, Zhao-Long Shen, Yuan Qian et al. · 0 citations
Preprint Sep 2026

Find Something You Can't Do: Agentic Real-World Reinforcement Learning for Self-Improving VLA Models

FIND is introduced, an agentic real-world RL framework that closes the loop between scene understanding, weakness-aware practice, self-evaluation, and policy improvement in a persistent workspace and reframes autonomous practice as a scene-conditioned, performance-aware task-selection problem.

Yuan Fang, Ze-Chu Li, Hao-Lei Tong et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Fewer Steps, Better Actions: Rethinking Flow-Matching Inference for VLA Policies

Vision-language-action (VLA) policies based on flow matching generate action chunks through repeated evaluations of an action expert. Increasing the number of integration steps raises inference cost, but does not necessarily improve closed-loop success. We propose Coda, which reallocates part of this integration budget...

Zhi-Peng Tang, Xin-Da Chen, Wei-Ning Rao et al. · 0 citations
#machine learning Preprint Sep 2026

Reinforcement Learning for Real-Time Vision-Language-Action Policies

This work instantiates Real-Time EXPO-FT, an RL framework for finetuning real-time VLA policies that meets the real-time control requirements of dynamic real-world manipulation, demonstrating rapid, sample-efficient adaptation to challenging real-world dynamics.

Perry Dong, Kuo-Han Hung, D. Sadigh et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.