Skip to content

Author

Zeyuan Liu

3 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Privy to the Foil: Recasting Value Estimation with a Self-Privileged Critic for RLVR

Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse terminal rewards, and actor-critic methods such as PPO address this by learning value functions to construct token-level advantages. Their effectiveness, however, hinges...

Kun Liang, Chenming Tang, Clive Bai et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Diffusion Reward Models

Reward models underpin the alignment of large language models, yet the dominant designs reduce each prompt--response pair to a point estimate or to a distribution from a fixed parametric family. This is at odds with human preference, which is inherently multimodal: the same response can be reasonably judged in many way...

Xiang-Yang Wang, Bing-Xiang He, Ze-Yuan Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability?

StudyBench is introduced, a controlled physics benchmark that directly measures how efficiently a self-evolution method converts training material into capability, and turns self-evolution progress from an open-ended pursuit into a measurable target for future research.

Ying-Hao Chen, Zi-Xi Chen, Bingxiang He et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.