Curriculum Learning for Countdown Reasoning in RL Fine-Tuning: Static Schedules Help, Adaptive Frontiers Forget
The novelty is twofold: tying curriculum design to the non-degeneracy of the RLOO advantage, and a zero-overhead adaptive sampler that discovers difficulty from reward rather than from a hand-set proxy.
Fengzhou Li
· 0 citations