A survey of reward hacking in agentic large language model systems
Large language models (LLMs) deployed as agentic systems capable of tool use, code execution, file manipulation, and multi-step planning inherit and amplify the classical reinforcement learning problem of reward hacking. This survey synthesizes how proxy-based alignment and evaluation failures manifest across modern LL...