Skip to content

Author

Bingxiang He

We have 9 of 35 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations
Preprint Aug 2026

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchm...

Yuhao Zhan, Bingxiang He, Zecong Tang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Rethinking On-Policy Distillation of Large Language Models II: One Training Example

On-policy distillation (OPD) combines student-generated rollouts with dense token-level supervision from a teacher. Existing work has mainly studied its algorithmic behavior, leaving the role of training data unclear. We examine this role at the data-minimal limit by training on a single query. One-shot OPD keeps impro...

Zi-Xuan Fu, Bing-Xiang He, Yu-Xin Zuo et al. · 12 citations · ⚡1
#artificial intelligence Preprint Sep 2026

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

StudyBench is introduced, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability, and turns self-evolution progress from an open-ended pursuit into a measurable target for future research.

Ying-Hao Chen, Zi-Xi Chen, Bingxiang He et al. · 0 citations
Preprint Aug 2026

AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement

This work presents AI4AI\mbox{-}Bench, 10 frozen research repositories spanning 10 training algorithm families, and releases the task suite, the evaluators and every scored submission, so that the measurement can be repeated as these systems change.

Yi-Zhe Chi, Wenyi Li, De-Yao Hong et al. · 5 citations
Jul 2026

Weak-to-Strong Generalization via Direct On-Policy Distillation

Direct On-Policy Distillation (Direct-OPD) is proposed, which transfers the teacher's RL-induced policy shift instead of running sparse-reward RL on the target model and consistently leverages weaker teachers to improve stronger target models.

Shiyuan Feng, Huan Gao, Haohan Chi et al. · 8 citations · ⚡1
#artificial intelligence Preprint Aug 2026

On-policy Distillation with Verifiable Reward

This work reformulates the implicit reward of sampled-token OPD based on trajectory correctness, then applies a ReLU gating mechanism to ensure that correct trajectories receive non-negative rewards and incorrect ones receive non-positive rewards, making it readily combinable with any policy gradient algorithm, such as...

Wen-Ze Lin, Jia-Le Zhao, Xi-Tai Jiang et al. · 7 citations · ⚡4
#software testing Preprint Aug 2026

SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration?

SWE Refactor Bench is introduced, a benchmark comprising 20 whole-repository migrations, covering 4 kinds of technical debt, and SWE Refactor Bench is positioned as a rigorous testbed for developing coding agents for reliable whole-repository migrations.

De-Yao Hong, Yi-Zhe Chi, Wen-Yi Li et al. · 3 citations · ⚡1
#artificial intelligence Preprint Aug 2026

Open-MOPD: Diagnosing and Fixing Capability Imbalance in Multi-Teacher On-Policy Distillation

Open-MOPD, a principled framework incorporating token-share balancing, gap-aware dynamic budget allocation, and student reward refresh, systematically restore cross-domain balance, elevating headroom recovery from 35.6% to 83.4% in a single deployable student.

Huan Gao, Haohan Chi, Yong Yan et al. · 4 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.