Preprint
Aug 2026
Agentic Reinforcement Learning with Observation-Calibrated Self-Distillation
Observation-Calibrated Self-Distillation (OCSD), which contrasts two structurally matched replay views, Full and Observation-Ablated, to derive an observation residual that discounts score changes shared by the replay scaffold, and applies this residual to modulate token-level GRPO updates at high-uncertainty steps, while preserving the trajectory-level update direction.
Y. Yang, Congming Qin, Xiaodan Liu et al.
· 1 citation