ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
This work proposes ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization, enabling sustained performance improvements over extended training horizons.