Preprint
Aug 2026
Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization
This work trains language models with GRPO on multiple-choice math problems where the correct answer is always option A, then evaluates on an unseen test set with unbiased answer positions to find reasoning-answer decoupling, which separates capability loss from a learned, transferable shortcut.
Suyash Maniyar, Armaan Sandhu, Abhishek Mishra
· 0 citations