Adaptive Difficulty-Aware Curriculum Learning for RLOO
This work proposes three curriculum variants on top of a RLOO fine-tuning baseline that can concentrate RLOO training on this frontier and improve final performance on the Countdown arithmetic reasoning task.
Darren Chan, Jayna Huang, Sophie Zhang
· 0 citations