Preprint
Jul 2026
Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
VarIance Guided Online Rollout allocation (VIGOR) is proposed which instead of allocating a fixed rollout budget per example, begins with a small number of rollouts for all examples in a batch and iteratively allocates additional rollouts to those with the highest group reward variance until a fixed total rollout budget is reached.
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman
· 1 citation