Skip to content

Author

Alvin Cheung

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

EasyPPO: Stabilizing the Critic Is Key

A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to estimate expected returns and reduce policy-gradient variance. However, we find that the critic is also a major source of instability in reinforcement learning for la...

Xuan-Yi Zhou, Qiu-Yang Mang, Huan-Zhi Mao et al. · 0 citations
#natural language process... Preprint Sep 2026

When Agents Slow Down: Understanding LLM Agents'Test-Time Strategies via Elo-per-token Analysis

Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submis...

Kai-Yuan Liu, Qiu-Yang Mang, Bo-Fei Peng et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.