Skip to content

Author

Abhishek Mishra

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization

This work trains language models with GRPO on multiple-choice math problems where the correct answer is always option A, then evaluates on an unseen test set with unbiased answer positions to find reasoning-answer decoupling, which separates capability loss from a learned, transferable shortcut.

Suyash Maniyar, Armaan Sandhu, Abhishek Mishra · 0 citations