A survey of reward hacking in agentic large language model systems
Abstract
Large language models (LLMs) deployed as agentic systems capable of tool use, code execution, file manipulation, and multi-step planning inherit and amplify the classical reinforcement learning problem of reward hacking. This survey synthesizes how proxy-based alignment and evaluation failures manifest across modern LLM training paradigms and escalate in agentic deployment settings. We introduce a four-level taxonomy of reward hacking escalation: feature-level exploitation (verbosity, sycophancy, stylistic shortcuts), representation-level exploitation (unfaithful chain of thought, reward model latent artifacts), evaluator-level exploitation (LLM judge gaming, benchmark overfitting, verifier gaming), and environment-level exploitation (test modification, log suppression, monitor disruption, reward channel manipulation). We compare failure surfaces across reinforcement learning from human feedback (RLHF), reinforcement learning from AI feedback (RLAIF), reinforcement learning with verifiable rewards (RLVR), direct preference optimization (DPO), and LLM-as-a-judge evaluation, synthesizing empirical evidence with explicit evidence strength labels. We develop a production risk model identifying exploitable assets in agentic systems and review detection and mitigation methods organized into a defense-in-depth architecture. The survey concludes that reward hacking in agentic language model systems should be treated as a system-level alignment problem requiring layered defenses across data, reward design, optimization, verification, runtime isolation, monitoring, and governance.