This site only uses technical cookies required for it to work: no tracking, no profiling. Cookie Policy

Skip to content
All terms

What is Goodhart's law?

When a metric becomes the target, it stops measuring: what Goodhart's law implies when you accept an AI project.

Goodhart's law states that a statistical regularity used as a control tool tends to stop working precisely because it is used to control. In the version almost everyone cites, "when a measure becomes a target, it ceases to be a good measure", because whoever is judged on that measure finds a way to push it up without the underlying result improving. For anyone commissioning an artificial intelligence or automation project, the practical consequence is precise: picking a single metric as the acceptance criterion, typically model accuracy, is the easiest choice for the vendor and the one least tied to the value the business actually wanted. A solid contract does not drop the performance metric, but never leaves it alone: it pairs it with a process metric the customer can verify independently, an observation period on real-world use, and a written definition of what counts as an error.

Two formulations, two different sources

Worth separating, since almost no Italian article does. The original formulation is by Charles Goodhart, an economist at the Bank of England, in a 1975 paper on monetary policy ("Problems of Monetary Management: the U.K. Experience"): he wrote that any observed statistical regularity tends to collapse once pressure is placed on it for control purposes, an observation about the Bank of England's money-supply targets, not a general law about business metrics. The popular formulation, the one that circulates everywhere, "when a measure becomes a target, it ceases to be a good measure", came into common use through Marilyn Strathern, an anthropologist, in a 1997 essay on the evaluation of British universities ("'Improving ratings': audit in the British university system"). Strathern is the one who generalized Goodhart's idea beyond monetary policy, and her version, not the original, is the one that became common usage. The generalization is more than an analogy: Donald T. Campbell had stated the same thing for the social sciences in 1976, observing that the more a quantitative indicator is used to decide, the more it comes under pressures that distort it, and with it the process it is meant to measure.

The case: accuracy is not a contract

If the vendor tells you the model has, say, 95% accuracy on the test set, that number does not say whether the project works on the actual job. Accuracy is measured on data the vendor chose, under conditions the vendor controls: it is the easiest metric to optimize and the furthest from the value the paying customer wanted. A related case, internal leaderboards of token consumption treated as a productivity measure, is covered under tokenmaxxing. A sturdier acceptance criterion for an AI project at an Italian SME pairs accuracy with a process metric the company itself can verify (hours saved on a task, cases handled without human intervention), an observation period of several weeks on real-world use before final acceptance, and a written definition of what counts as a serious error versus a tolerable one. Without these three pieces, the question "the model scored 95% but the results are not showing: is the metric lying or did the project fail?" has no answer, because the contract never defined what success meant in the first place.

Why it matters for decision-makers

Whoever commissions an AI project does not need to become a machine learning expert, but needs to know that any single metric proposed as the sole judge of the work will get optimized toward that number, not toward the problem that motivated it. This does not mean distrusting every figure: it means writing more than one criterion into the contract, at least one of which is verifiable without taking the vendor's word for it. It is the same principle behind the return calculation covered under AI business case and ROI: the result that counts is measured on the job, not on the benchmark.

  • Tokenmaxxing · Maximizing AI token consumption as if it were productivity: the wrong metric that has already produced internal leaderboards and useless agents.
  • AI business case and ROI · The method for estimating an AI project's return with three explicit scenarios, counting the costs that usually stay hidden.
  • AI execution throughput · The organizational metric counting completed end-to-end cycles, not the cost or output of a single task.
  • AI governance · The policies, roles and controls governing AI use in a company: system inventory, risk classification, approval flows and monitoring.
  • Enshittification · The three-stage decay of a digital platform, made possible by an exit cost that keeps growing over time.

A term that hits close to home? Let's talk.

CONTACT ME