Skip to content
Review Open access

A survey of reward hacking in agentic large language model systems

Aug 2026 · Discover Artificial Intelligence · Vol 6 · 0 citations · 81 references
Computer Science

Abstract

Large language models (LLMs) deployed as agentic systems capable of tool use, code execution, file manipulation, and multi-step planning inherit and amplify the classical reinforcement learning problem of reward hacking. This survey synthesizes how proxy-based alignment and evaluation failures manifest across modern LLM training paradigms and escalate in agentic deployment settings. We introduce a four-level taxonomy of reward hacking escalation: feature-level exploitation (verbosity, sycophancy, stylistic shortcuts), representation-level exploitation (unfaithful chain of thought, reward model latent artifacts), evaluator-level exploitation (LLM judge gaming, benchmark overfitting, verifier gaming), and environment-level exploitation (test modification, log suppression, monitor disruption, reward channel manipulation). We compare failure surfaces across reinforcement learning from human feedback (RLHF), reinforcement learning from AI feedback (RLAIF), reinforcement learning with verifiable rewards (RLVR), direct preference optimization (DPO), and LLM-as-a-judge evaluation, synthesizing empirical evidence with explicit evidence strength labels. We develop a production risk model identifying exploitable assets in agentic systems and review detection and mitigation methods organized into a defense-in-depth architecture. The survey concludes that reward hacking in agentic language model systems should be treated as a system-level alignment problem requiring layered defenses across data, reward design, optimization, verification, runtime isolation, monitoring, and governance.

Read PDF

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.