Enhancing RLOO with Dense Symbolic Rewards and Bandit-Driven Curricula
This project explores various reinforcement learning techniques to improve model arithmetic reasoning beyond supervised imitation, and evaluates two RLOO extensions: dense rewards, which credits valid intermediate arithmetic steps, and an adaptive curriculum that uses a bandit policy to sample problem difficulty buckets based on recent learning signal.
Rinnara Sangpisit, Irawadee Thawornbut
· 0 citations