Skip to content
Preprint

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

Aug 2026 · 0 citations · 30 references
Computer Science

TL;DR

A moment-based perspective on policy optimization for LLM reasoning is introduced by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments, leading to MMPO, a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution.

Abstract

Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-based perspective on policy optimization for LLM reasoning by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments. Under this perspective, many existing methods optimize only a single moment of the failure-probability distribution, leaving its broader distributional structure largely uncharacterized. We propose \textbf{M}ulti-\textbf{M}oment \textbf{P}olicy \textbf{O}ptimization (MMPO), a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution. MMPO admits a direct operational interpretation as minimizing the expected truncated time required to obtain the first successful response. Beyond MMPO, we further develop a general moment-transformation framework that systematically induces different moment profiles and provides a unified view of a broader family of policy optimization objectives. Experiments across five mathematical reasoning benchmarks and models of different scales demonstrate that MMPO consistently outperforms strong baselines. We hope this moment-based perspective offers new insights into the design of policy optimization objectives for LLM reasoning.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

Information-Time Proximal Policy Optimization

RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, whi...

Yong-Cheng Zeng, Xin-Yu Cui, Yan Song et al. · 0 citations
#artificial intelligence Preprint Sep 2026

GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the...

Kai-Chen Zhang, Yuzhong Hong, Jun-Wei Bao et al. · 0 citations
Preprint Aug 2026

Learning When to Think: Adaptive Reasoning for Test-Time Compute Allocation

A reasoning model is built that adaptively chooses how much to reason for each problem, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random.

Gijs Kassenaar, Zhao Yang, Vincent François-Lavet · 1 citation
#artificial intelligence Preprint Sep 2026

MaD-RL: Matching Distributions for Calibrating LLMs with Reinforcement Learning

This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...

Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al. · 0 citations
Conference Aug 2026

Post-Training for Reasoning LLMs with Reinforcement Learning: A Stability–Efficiency Perspective

Reinforcement Learning (RL)-based post-training has become a key mechanism for improving reasoning Large Language Models (LLMs), especially when automated verifiers enable Reinforcement Learning with Verifiable Reward (RLVR). Recent methods, however, are increasingly modular: many apparent algorithmic advances recombin...

Liu Yang, Han Zhu, Zheng-Yang Zhong et al. · 2 citations
Preprint Aug 2026

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

Group Planning-aware Policy Optimization (PlanPO) is proposed, a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns that enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generat...

D. Liang, Liyuan He, Xuan Feng et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.