SFT-Estimated Curriculum Learning for Rule-Based RLOO Fine-Tuning
This project studies whether curriculum-based prompt ordering can make RL fine-tuning more stable and sample-efficient for language-model reasoning, and when curriculum structure helps, when it fails, and what failure modes appear in small-scale online RL fine-tuning.
Vanessa Felix
· 0 citations