Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...
Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al.· 0 citations
Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics...
Xiao-An Xu, Si-Yuan Liu, Shuo Wang et al.· 0 citations
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchm...
Yuhao Zhan, Bingxiang He, Zecong Tang et al.· 0 citations
On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps impro...
StudyBench is introduced, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability, and turns self-evolution progress from an open-ended pursuit into a measurable target for future research.
Ying-Hao Chen, Zi-Xi Chen, Bingxiang He et al.· 0 citations
REFACT is an adaptive fact-restatement citation framework that enables LLMs to determine when contextual grounding is needed and selectively restate source facts at appropriate levels of detail for reliable reasoning.
Zhensheng Jin, Xin Dai, Zhenghao Liu et al.· arXiv.org· 0 citations
Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.
Zhu Zhang, Ji-Xun Wang, Xiao-An Xu et al.· 2 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.