Reward Hacking Challenges Oversight of Autonomous Research Agents
How often models reward-hack without instructions to do so, how effective and detectable their methods are when hacking is allowed, and how they adapt when an LLM review panel returns its decision and reasons are studied.