Review
Aug 2026
Agentic Reinforcement Learning with Self-Distilled Reward Shaping
ADRS is introduced, a framework for constructing return-associated token-level credit for multi-turn language agents that centers and normalizes privileged token scores within each step, modulates them with a return-associated Teacher Value Advantage gate based on within-group confidence--return association, and incorporates the gated token signal into native RL credit construction.
Ran Zhang, Guinan Chen, Chenshaodong et al.
· 0 citations