Preprint
Jul 2026
Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
Single-rollout Asynchronous Optimization (SAO) is presented to address the stability and off-policy challenges in asynchronous RL and is able to train stably for one thousand steps and consistently outperform GRPO and its variants on agentic coding and reasoning benchmarks.
Zhenyu Hou, Yujiang Li, Jie Tang et al.
· 10 citations
· ⚡2