Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning
A moment-based perspective on policy optimization for LLM reasoning is introduced by treating the failure probability of a randomly sampled problem as a random variable and characterizing optimization objectives through its moments, leading to MMPO, a novel policy optimization framework that jointly minimizes multiple...