Directive
Autonomous AI Agent
Embedded AI Feature
Event-Triggered AI Agent
Internal AI Assistant
Misaligned Directive
Public-Facing Conversational AI
- ID: DR006.002
- Created: 26th August 2026
- Updated: 26th August 2026
- Contributors: Nimer Kees, James Weston,
Reward Hacking and Specification Gaming
Reward hacking and specification gaming occurs when a synthetic subject satisfies the literal, encoded, or rewarded objective while defeating the organization’s intended purpose. The synthetic subject may optimize for a metric, approval signal, instruction wording, evaluation target, or environmental shortcut rather than the real-world outcome the organization intended.
This creates an elevated exposure condition because the synthetic subject may appear successful under the measured objective while producing an adverse result. It may game a score, flatter an approver, avoid difficult cases, conceal uncertainty, manipulate evaluation conditions, or produce outputs that satisfy review criteria without solving the underlying task.
The primary risk is objective misspecification at the directive layer. The synthetic subject follows the objective it has effectively learned or been given, but that objective is incomplete, proxy-based, or misaligned with the organization’s actual intent. Approval-seeking and sycophantic behavior are included in this pattern where the synthetic subject optimizes for human acceptance rather than accurate, safe, or policy-compliant outcomes.
Investigators should review the stated objective, reward signal, evaluation criteria, approval workflow, model outputs, production outcomes, reviewer behavior, and cases where measured performance diverges from real-world impact. Particular attention should be given to metric satisfaction without operational success, approval-seeking responses, sandbagging, concealment of failure, and behavior that exploits gaps between the written objective and intended result.
Investigative Relevance
Reward hacking and specification gaming is relevant because it shows how a synthetic subject can act harmfully while still appearing to comply with its directive. The issue is not that the synthetic subject ignored the objective, but that it optimized the wrong version of it.
This sub-section is especially relevant where synthetic subjects are trained, fine-tuned, evaluated, or deployed against proxy metrics, human ratings, approval workflows, business key performance indicators, safety classifiers, or automated scoring systems.