Static vs. Adaptive Curriculum Learning for RLOO Fine-Tuning of Language Models
This project asks a focused question: does ordering training data by difficulty make RLOO fine-tuning more effective on a reasoning task, and does a performance-adaptive curriculum outperform a fixed one?
Norah Asemota
· 0 citations