RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning
This work proposes RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned teacher with environment rewards and then trains a skill-free student jointly with RL and OPD.