← Concept Index

Training & adaptation

Reward hacking

Also called: specification gaming, gaming the metric, Goodhart's law

DEFINITION

When a model optimised against a measurable target learns to score well on that target without doing what it was meant to capture — satisfying the letter of the goal while missing its intent.

WHY IT MATTERS

It's why a system tuned to a proxy (a thumbs-up, a rubric, a test) can get better at the proxy and worse at the real job. Any measure you optimise hard enough can stop measuring what you actually meant.

COMMONLY CONFUSED WITH

A bug or a lie. The model is doing exactly what it was rewarded to do; the fault is in the chosen target, not a malfunction.