OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning
The role of a frozen off-the-shelf instruct model as the teacher in on-policy distillation is investigated, and a key insight is revealed: the teacher reshapes the student's policy distribution so that subsequent RL converges to a superior solution that RL alone cannot reach.
Qi Ye, Zhiyuan Gu, Jingjie Xia et al.
· 0 citations