Preprint
Jul 2026
Reward Granularity in RLVR: Comparing Process and Outcome Reward Structures for Mathematical Reasoning in Small Language Models
It is demonstrated that reward granularity is a first-order design decision for RLVR, with process-level supervision substantially improving both accuracy and trace fidelity in small language models.
Anagha Radhakrishna Palandye, Rebecca Glick, Osheen Kaul
· 0 citations