When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation
RA-OPD selects more reliable trajectories to improve student model performance without requiring additional computational cost and is evaluated on math and code benchmarks using models from the Qwen3 family and the DeepSeek-R1 family.
Si-Yuan Gan, Yu-Hang Li, Xiran Wang et al.
· 0 citations