Skip to content

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning

Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advanta...

Mu-Yang Li, Jie Yang, Zheng-Yu Fang et al. · 1 citation · ⚡1
#machine learning Preprint Sep 2026

LastOPD: Taming Collapse in Latent On-Policy Distillation

LastOPD is proposed, which applies the latent signal only at the last-layer state, the common interface both LM heads read, and only during a 10-step crossfade into token-level OPD, which keeps the useful part of the latent signal and hands the student to token-level supervision before the collapse sets in.

Jie Yang, Zheng-Yu Fang, Ze-Lin Xu et al. · 2 citations · ⚡1

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.