Advantage-Driven Synthetic Curriculum for Reinforcement Learning based Fine-Tuning of Large Language Models
ADSC is introduced, a curriculum learning layer for REINFORCE Leave-One-Out that uses a signal already computed by RLOO to identify prompts near the student’s current learning frontier and uses a multi-armed bandit router to sample from these difficulty buckets.
Nathaniel Demchak, Pravin Ravishanker, Oscar Li
· 0 citations