A call centre's average handling time sat near six minutes for years, an honest number nobody paid much attention to. Then leadership made it a target: five minutes, a quarterly bonus attached. Handling time fell to four and a half within two quarters. Problem resolution fell with it. Reps rushed callers off the line to protect the number, callbacks and complaints both climbed, and the dashboard read green the entire time.
That's Goodhart's Law, and it's a different failure from the one most measurement advice warns about.
Goodhart's Law, in one line: when a measure becomes a target, it stops being a good measure. Once people know a number is being watched and rewarded, they optimise the number, not the thing it was standing in for.
A different problem from picking the wrong metric
Most advice about bad measurement is really about vanity metrics: a number that was never meaningful to begin with. Follower count, page views, email opens. Those fail on day one, because they never told you anything true.
Goodhart's Law is the harder problem. It happens to a metric that was meaningful, one that genuinely tracked something real, corrupted specifically because it became consequential. Average handling time really did measure service quality, until the bonus made it worth gaming. The metric didn't start broken. Making it a target broke it.
Five ways it shows up
The pattern repeats across sectors, and it always has the same shape: the target gets hit, and the thing the target was supposed to protect gets worse.
| Where | The metric | The target got hit | What actually happened |
|---|---|---|---|
| Call centre | Average handling time | Time fell | Problem resolution fell too; reps rushed callers off the line |
| University | Graduation rate | Rate rose | Admission standards dropped; easier intake, not better teaching |
| Hospital | Waiting-time compliance | Target met | Patients reclassified between queues, not actually seen faster |
| Service desk | Tickets closed | Closure rate rose | Recurring problems went unresolved, just reopened as new tickets |
| Project delivery | On time, on budget | Both hit | Benefits were never realised; delivery succeeded, value didn't move |
Every row on the right is a dashboard that stayed green. Nobody had to hide anything. The number simply stopped meaning what it used to mean, and the people closest to it were the last to notice, because from where they sat, they were succeeding.
The same mechanism as Measurement Theatre
Measurement Theatre is what this looks like at the organisational scale: the dashboard immaculate, the board reassured, and nobody actually better off while the work the measurement was meant to protect quietly stops happening. Goodhart's Law is the mechanism underneath it, not a separate failure. It's the specific way a single metric gets there: not through fraud, through people rationally responding to what they're being measured on.
How to catch it before the dashboard tells you
A metric that's just being watched is safer than one that's consequential. The moment a number carries a bonus, a promotion, or a public target, ask three questions.
Has this number stopped moving in ways it used to? Real underlying variation should still show up somewhere. A metric that's suspiciously smooth, always just past the target, rarely below it, is often a metric being managed rather than a process being improved.
Does it still correlate with the outcome it was supposed to stand for? Average handling time used to track service quality. Once the correlation breaks, that's the tell, not a warning sign to watch for later.
Is there a second check that isn't itself a target? The fix isn't abandoning the metric. It's holding two things at once: the number people are asked to hit, and a harder-to-game outcome check that nobody's bonus depends on. OKRs vs KPIs covers the same discipline from the other direction, holding a steady-state number and a real bet on the same team without letting either one substitute for the other.
The same pattern, now showing up in AI
Goodhart's Law didn't stay confined to call centres and hospitals. The same mechanism is starting to show up wherever an AI model gets optimised against a benchmark: the benchmark score climbs, and the capability it was supposed to measure doesn't move by the same amount, because the model learned the benchmark, not the underlying skill it stands in for. Same three questions above, applied to a newer kind of target.
This is Goodhart's Law seen through Measurement Theatre's own argument: a dashboard can be entirely honest and still stop meaning anything, the moment the number on it becomes the goal. The same failure pattern, in different costumes, runs through the phase every management system skips. The Shaping What Gets Built course teaches the oversight discipline that catches it before the target does the damage.
