This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignme...
Tian-Ze Xu, Yan-Zhao Zheng, Zhen-Tao Zhang et al.· 0 citations
CC-OPD (Counterfactual Constraint-Conditioned On-Policy Distillation), which inverts the standard supervision-generation direction in distillation, and achieves the highest average among all evaluated student-training methods.
Yan-Zhao Zheng, Yuan-Qiang Yu, Tian-Ze Xu et al.· 0 citations
PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance.
Yuan-Qiang Yu, Yan-Zhao Zheng, Zhen-Tao Zhang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.