Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage for RL and ask whether its benefits extend beyond improvements in the distilled model's initial accuracy. Under shared RL settings, students...
Shuai Dong, Yong-Fu Zhu, Yu-Qi Xu et al.· 0 citations
On-policy distillation, in which a teacher corrects samples that the student itself generates, presupposes that the two models speak the same language: identical VAE latents, matching architectures, and a common timestep grid. We ask what happens when none of this holds, as when the strongest teacher available and the...
This paper presents Poly-OPD, a framework that can consolidate complementary strengths of heterogeneous teachers into a single compact flow-matching student, and uses a gradient compatibility diagnostic to organize its adapters.
Siming Fu, Haojun Xu, R.Z. He et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.