Jun 2026
Behavior Cloning is Not All You Need: The Optimality of On-Policy Distillation for Noisy Expert Feedback
A noisy expert model is proposed to explain why OPD can outperform SFT when training language models from imperfect teachers, and it is proved that online interaction with the noisy expert via a novel variant of OPD enables polynomial dependence on the horizon in general.
V. Sriraman, Peihan Liu, Daniel Hsu et al.
· arXiv.org · 0 citations