D-MOPD (Dynamic Domain ScheDuling for MOPD), a zero-overhead scheduler that repurposes the per-domain reverse-KL signal already produced during training to adapt the domain mixture online to adapt the domain mixture online.
Zechen Sun, Zhi-Wei Zhang, Fei Zhao et al.· 1 citation
CSM-MTBench is introduced, a benchmark covering five Chinese-foreign language directions and consisting of two expert-curated subsets: Fun Posts, featuring context-rich, slang- and neologism-heavy content, and Social Snippets, emphasizing concise, emotion- and style- driven expressions.
On-policy distillation (OPD) accelerates post-training by providing dense token-level supervision from a frozen teacher on the student's own rollouts. Vanilla OPD applies this supervision uniformly across prompts, without checking whether the teacher is reliable for each prompt. Because reverse KL is mode-seeking, a co...
Zhi-Wei Zhang, Zechen Sun, Fei Zhao et al.· 5 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.