Preprint
Aug 2026
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs
The sample efficiency and scalability of RL post-training for video MLLMs and introduces OraRL, a decoupled advantage estimator that scales with model size and data, surpassing its backbone from 0.8B to 9B and GRPO up to 100k prompts.
Yunheng Li, Guohong Mu, Hao Li et al.
· 0 citations