Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from"Rollout Silencing"and low-quality gradient signals in standard sampling procedures. In this work, we propose...