ANCHOR: Taming Entropy Dynamics for Stable and Efficient Reasoning of Large Language Models
It is shown that policy entropy bounds both the policy gradient and probability update norms; consequently, entropy collapse effectively stops reward signal backpropagation, preventing further policy learning regardless of data quality.