Programmatic Reasoning for Countdown: Learning to Generate Executable Python-Style Verifications
This work asks whether replacing free-form reasoning with a restricted, executable program turns this sparse reward into a graded, more interpretable training signal, using a standard SFT → RLOO pipeline as the reference point.
O. Ivankiv, Henry Zhou
· 0 citations