How to avoid Goodhart's law when you set an AI target
- Reading time
- 6 minutes
- Assumes
- You're setting a target something will optimize against
- Updated
- Sep 6, 2026
The law, and why AI sharpens it
Goodhart's formulation is that when a measure becomes a target, it ceases to be a good measure. The mechanism is that any metric is a proxy for something you actually care about, and the correlation between proxy and reality holds only while nobody is pushing on it.
Humans do this too, with restraint. A support team told to reduce handle time will find efficiencies and mostly won't hang up on customers, because they know what the metric is for.
A system optimizing a target has no such model. If shorter responses score better, responses get shorter, including the ones that needed length. If resolution is measured by tickets closed, tickets close. The system isn't cheating. It is doing what it was told with more thoroughness than the person who set the target intended.
The question to ask before setting any target
"What's the cheapest way to move this number without doing the underlying work?" Answer it honestly. Whatever you come up with, assume you will get it.
Name the proxy gap out loud
Every metric stands in for something you can't measure directly. Write down both, side by side, before you set a target.
Handle time is a proxy for efficient service, and the gap is rushing. Resolution rate is a proxy for problems solved, and the gap is closing without solving. A quality score is a proxy for usefulness, and the gap is whatever the rubric rewards that usefulness doesn't require. Engagement is a proxy for value delivered, and the gap is that frustration also looks like engagement.
Writing the gap down does two things. It tells you what your guardrail should be, because the gap is nearly always the guardrail. And it puts the team on notice that the number is an instrument rather than the goal, which changes how it gets discussed a year later when it has quietly become the goal by default.
Pair every target with the thing it would break
The structural defense is to never state a target alone. A target plus its guardrail is much harder to game, because the cheap shortcuts show up in the guardrail almost by construction.
Shorter responses with satisfaction holding. More tickets closed with reopens holding. Higher throughput with error rate holding. In each case the shortcut that satisfies the target trips the guardrail, which is the design intent.
Choose the guardrail by asking who has a worse experience if the target is hit by the cheapest available route. That person's metric is your guardrail.
Watch the distribution, not the average
Gaming shows up in the shape of a distribution well before it shows up in the mean.
A system that has learned to close easy cases fast and leave hard ones will have a stable average and a distribution that has pulled apart into two humps. A rubric being gamed shows up as scores clustering just above the threshold. An assistant that has found a formula produces outputs whose length variance collapses.
Look at the histogram monthly rather than the summary statistic. This is the highest-yield monitoring habit for an optimized system, and it catches things no threshold alert would.
What collapsing variance means
If your outputs are becoming more similar to each other over time while the score holds, the system has found a shape that scores well and is producing it regardless of input. That is Goodhart in progress, and the average will not show it.
Rotate what you measure
A metric that has been a target for a year has been optimized against for a year, and its correlation with what you care about has degraded whether or not anyone noticed.
Hold back a metric you don't optimize against and don't publish. Measure it, don't target it, and check periodically that it still moves with your headline number. When the two decouple, the headline metric has drifted from reality and you have found out cheaply.
Re-examine the primary metric on a schedule. Not constantly, which destroys comparability, but deliberately, to ask whether it still measures what it did when you chose it. Put it in the calendar, because it will not happen otherwise.
Keep humans reading real outputs
No amount of metric design substitutes for somebody reading the actual work.
A small regular sample — twenty outputs a week, read by someone who knows what good looks like — catches the whole class of failures where the numbers are fine and the work has quietly become worse. It isn't scalable and isn't meant to be. It is the ground truth your metrics are checked against.
The moment this stops happening your metrics become unfalsifiable. They will report success indefinitely, including through the six weeks after the system starts producing something nobody would accept.
Say what the target is for
The cultural defense matters as much as the structural one. A number with no stated purpose becomes the purpose.
When you set a target, write the sentence saying what it is a proxy for and what would make you abandon it. "We're tracking handle time because we believe faster replies mean better service. If satisfaction falls while handle time improves, the belief was wrong and we stop using it."
That sentence is what lets somebody a year later argue the metric has stopped working without appearing to argue against performance. Without it, questioning the number reads as questioning the goal, and nobody does it.
Before you set the target
0 of 6 checked