Preprint
Aug 2026
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
It is found that training to reason in the native language often leaves only a small gap to training for English reasoning, and RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.
Konstantin Dobler, Federico Scozzafava, Jonathan Janke et al.
· 0 citations