Skip to content

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

Sep 2026 · 1 citation · 36 references
Computer Science

TL;DR

Value Flattening is identified as an important yet overlooked failure mode of critic learning in standard PPO and a simple sparse supervision strategy can mitigate it; SParse Proximal Policy Optimization is introduced, which applies the value loss to only a few well-separated states in each response to mitigate both effects.

Abstract

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empirical analyses relate Value Flattening to an implicit variance penalty in the critic loss and redundant updates from temporally correlated states with similar gradients. Motivated by these findings, we introduce SParse Proximal Policy Optimization (SP$^3$O), which applies the value loss to only a few well-separated states in each response to mitigate both effects. Experiments on Qwen3-Base show that SP$^3$O with only three states supervised per response can mitigate Value Flattening and consistently improve the learned policy across model sizes and evaluation suites. Together, our results identify Value Flattening as an important yet overlooked failure mode of critic learning in standard PPO and show that a simple sparse supervision strategy can mitigate it.

View source

Similar papers

#machine learning Preprint Sep 2026

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors simultaneously, making the contribution of rollout policy difficult to isolate. W...

Julianna Piskorz, Antonin Berthon, M. van der Schaar · 0 citations
Preprint Aug 2026

Start Classifying: Categorical Critics for LLM Reinforcement Learning

Proximal Policy Optimization (PPO) for large language models typically trains its critic by mean-squared-error (MSE) regression on scalar value targets. Although scalar MSE is statistically valid for estimating the conditional expected return, sparse binary rewards in reinforcement learning with verifiable rewards (RLV...

Zhi-Jian Zhou, Long Li, Xuan Zhang et al. · 2 citations · ⚡2
#small language model Preprint Aug 2026

Boosting LLM Exploration via Weak-Model Guidance in RLVR

This work empirically study the potential of outer prefixes, revealing the mechanism of the impact of distributional discrepancy to the exploration dynamics in RLVR training and efficiently mitigates entropy collapse without requiring additional SFT, intricate reward designs, or complex prompting.

Xin Shen, Huishuai Zhang, Peng Li et al. · 0 citations
#artificial intelligence Preprint Aug 2026

BCPPO: Bachelier-Inspired Constrained Proximal Policy Optimization for Tail-Risk-Aware Safe Reinforcement Learning

BCPPO (Bachelier-Inspired Constrained Proximal Policy Optimization), a proximal policy optimization (PPO) method, supports a practical balance among reward, caution around cost predictions that vary across trained critics, and policy-only deployment.

Dong-Sheng Hou, Yanqiao Chen, Yu-Han Rui · 0 citations
Preprint Aug 2026

SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement Learning

This work proposes Single-rollout Autoregressive Policy Optimization (SAPO), a low-memory and compute-efficient framework in which the policy and value functions share a single autoregressive backbone, and introduces a trajectory-level generalized advantage estimator that combines lambda-returns with batch normalizatio...

D. Liang, Lang Feng, Bo An et al. · 1 citation
Preprint Aug 2026

ReBRAC-v2: The Return of the King

ReBRAC-v2 is introduced, which directly trains an exact-likelihood normalizing flow as the RL actor, combines likelihood, MSE, and MAE behavior regularization, and integrates a classification-based residual critic, staged optimization, and multi-sample test-time action selection, and ranks first in eight categories.

Denis Tarasov, Robert K. Katzschmann · 0 citations

Related blog posts

MIT News · Artificial Intelligence Sep 29, 2026

Who we become when we talk to machines

Professor Sherry Turkle’s new book, “Artificial Intimacy,” offers a withering critique of chatbots and the antisocial dynamics she believes they encourage.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.