Skip to content

Author

Chaojun Xiao

We have 7 of 75 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

A*-Thought-V2: Efficient Latent Reasoning via Geometric Dynamics of LLM

Chain-of-Thought (CoT) improves the reasoning ability of Large Language Models (LLMs) but incurs substantial computation and context costs. Existing methods either lose intermediate information through hard pruning or lack a principled criterion for continuous compression. We present A*-Thought-V2, a geometric dynamics...

Xiao-An Xu, Si-Yuan Liu, Shuo Wang et al. · 0 citations
Preprint Aug 2026

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchm...

Yuhao Zhan, Bingxiang He, Zecong Tang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps impro...

Zi-Xuan Fu, Bing-Xiang He, Yu-Xin Zuo et al. · 12 citations · ⚡1
#artificial intelligence Preprint Sep 2026

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

StudyBench is introduced, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability, and turns self-evolution progress from an open-ended pursuit into a measurable target for future research.

Ying-Hao Chen, Zi-Xi Chen, Bingxiang He et al. · 0 citations
Jul 2026

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

REFACT is an adaptive fact-restatement citation framework that enables LLMs to determine when contextual grounding is needed and selectively restate source facts at appropriate levels of detail for reliable reasoning.

Zhensheng Jin, Xin Dai, Zhenghao Liu et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Beyond Teacher Likelihood: Group-Calibrated On-Policy Distillation for Long-Context Reasoning

Group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance, and demonstrates that group-relative residual calibration can incorporate verifier outcomes without discarding dense token-level guidance.

Zhu Zhang, Ji-Xun Wang, Xiao-An Xu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.