Skip to content

Author

Baohua Dong

We have 5 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation

This work proposes ExpertAlign, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model, and demonstrates token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignme...

Tian-Ze Xu, Yan-Zhao Zheng, Zhen-Tao Zhang et al. · 0 citations
Book Open access Aug 2026

Distribution-Value Coevolution for Adaptive RLHF Data Scheduling

This work identifies and formalizes the Distribution-Value Coevolution principle: the training value of data is not intrinsic, but emerges dynamically from the interaction between data characteristics and the model's evolving capability boundary, and operationalizes this principle through a unified framework.

Zairun Yang, Yanbo Yang, Chenyi Zhou et al. · 0 citations
Book Open access Aug 2026

Distribution-Value Coevolution for Adaptive RLHF Data Scheduling

Reinforcement learning from human feedback (RLHF) has become the cornerstone of aligning large language models (LLMs) with human intent. Yet a fundamental question remains unaddressed: how should training data be scheduled when both the model's capabilities and the utility of data are constantly evolving? Current pipel...

Zairun Yang, Yanbo Yang, Chenyi Zhou et al. · 0 citations
#machine learning Preprint Aug 2026

PAC: Progress-Augmented Advantage Curriculum for Multi-Task Reinforcement Learning of LLMs

PAC, a Progress-Augmented Advantage Curriculum for multi-task RL of LLMs that combines two task-level signals: advantage-derived learnability, which measures the magnitude of the policy update a task can induce, and recent reward gains, which show whether those updates have improved task performance.

Yuan-Qiang Yu, Yan-Zhao Zheng, Zhen-Tao Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.