A moment-based perspective on policy optimization for LLM reasoning is introduced by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments, leading to MMPO, a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution.
Abstract
Reinforcement learning has become a central paradigm for improving the reasoning capabilities of large language models. Existing methods generally aim to reduce the failure probabilities induced across problems. In this paper, we introduce a moment-based perspective on policy optimization for LLM reasoning by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments. Under this perspective, many existing methods optimize only a single moment of the failure-probability distribution, leaving its broader distributional structure largely uncharacterized. We propose \textbf{M}ulti-\textbf{M}oment \textbf{P}olicy \textbf{O}ptimization (MMPO), a novel policy optimization framework that jointly minimizes multiple moments of the failure-probability distribution. MMPO admits a direct operational interpretation as minimizing the expected truncated time required to obtain the first successful response. Beyond MMPO, we further develop a general moment-transformation framework that systematically induces different moment profiles and provides a unified view of a broader family of policy optimization objectives. Experiments across five mathematical reasoning benchmarks and models of different scales demonstrate that MMPO consistently outperforms strong baselines. We hope this moment-based perspective offers new insights into the design of policy optimization objectives for LLM reasoning.
RLVR has substantially improved the reasoning capabilities of LLMs. However, existing methods typically parameterize temporal progression in the Markov Decision Process by token-by-token generation, despite the highly non-uniform information flow along autoregressive trajectories. In this paper, we propose InfoPPO, whi...
Yong-Cheng Zeng, Xin-Yu Cui, Yan Song et al.· 0 citations
Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the...
Kai-Chen Zhang, Yuzhong Hong, Jun-Wei Bao et al.· 0 citations
A reasoning model is built that adaptively chooses how much to reason for each problem, and the brief modes end up more accurate than \textsc{Long}, which shows that the router sorts problems by difficulty rather than at random.
Gijs Kassenaar, Zhao Yang, Vincent François-Lavet· 1 citation
This work proposes a general RL-based framework for Distribution Matching allowing matching the distribution of a latent categorical attribute of model outputs to a specified target distribution and proposes reward functions for other divergences such as KL and Jensen-Shannon and motivate them with theoretical justific...
Sourabh Kulkarni, Ksheeraj Sai Vepuri, B. Demir et al.· 0 citations
Reinforcement Learning (RL)-based post-training has become a key mechanism for improving reasoning Large Language Models (LLMs), especially when automated verifiers enable Reinforcement Learning with Verifiable Reward (RLVR). Recent methods, however, are increasingly modular: many apparent algorithmic advances recombin...
Liu Yang, Han Zhu, Zheng-Yang Zhong et al.· International Conference on...· 2 citations
Group Planning-aware Policy Optimization (PlanPO) is proposed, a simple yet effective RL method for learning generalizable planning abilities beyond task-specific high-quality behavior patterns that enables agents to actively learn generalizable and deliberate behaviors spanning interaction planning and textual generat...
D. Liang, Liyuan He, Xuan Feng et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.