Frontier Curriculum and Adaptive Test-Time Compute for Efficient RLOO
It is asked whether a single success-probability signal can allocate compute more efficiently in both phases on Countdown with Qwen2.5-0.5B.
Marco Vizcarra, Andy Kim
· 0 citations