Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Domain-Normalized MOPD is proposed, which keeps the routing and rescales each domain's feedback by its measured spread, and improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain.