Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...
Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al.· 0 citations
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps impro...
StudyBench is introduced, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability, and turns self-evolution progress from an open-ended pursuit into a measurable target for future research.
Ying-Hao Chen, Zi-Xi Chen, Bingxiang He et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.