Skip to content

Author

Irawadee Thawornbut

We have 1 of 1 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Enhancing RLOO with Dense Symbolic Rewards and Bandit-Driven Curricula

This project explores various reinforcement learning techniques to improve model arithmetic reasoning beyond supervised imitation, and evaluates two RLOO extensions: dense rewards, which credits valid intermediate arithmetic steps, and an adaptive curriculum that uses a bandit policy to sample problem difficulty buckets based on recent learning signal.

Rinnara Sangpisit, Irawadee Thawornbut · 0 citations