Preprint
Aug 2026
SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning
This work proposes Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone, and introduces a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalization.
D. Liang, Lang Feng, Bo An et al.
· 1 citation