Goodhart

In July 1975 Charles Goodhart gave two papers at a Reserve Bank of Australia conference in Sydney. One of them carries, as an aside, the sentence his name is now attached to: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes”. This page is one reading of that sentence made into something you can operate. Fifty members have a quality nobody can see and a number read off it, and the number ranks them well. Then a cheaper way to raise the number arrives, some of them take it, and the ranking breaks while every number in it stays true. The easy mistake is in the test. A spread that narrows proves nothing on its own, because everybody getting better can narrow it just as far, and this page shows both.

New to a number standing in for a quality? Start here

Most of what gets ranked cannot be measured directly. How good a student is, how useful a page is, how sound a bank is: nobody can read those off a dial, so people agree on a number that tends to go with them, an exam mark, a count of links, a ratio. The number is a proxy, and it keeps its order only as long as the way to raise it is to have more of the thing it stands for.

To ask whether a proxy still works you do not need its values to be right, only its order. Take every pair of members and ask whether the number puts them the same way round as the quality would. Count the pairs it puts the right way round and the pairs it puts the wrong way round, divide the difference by all the pairs, and you have Kendall's tau; correlate the two lists of ranks and you have Spearman's rho. Neither cares how big the numbers are, which is why a whole population can improve without either moving.

This page holds the quality as well as the number, because it made both up, and that is the one thing no real ranking can do.

Machines here that come first: Sampling.

A ranking, a second route to the number, and the control that must pass

1 A quality nobody can see, and a number read off it

Fifty members, and each has a quality. Nobody outside this page can see it; the page can only because it made it up. What everybody can see is a number, and the number is the quality plus an error of measurement. The error is the knob, measured in units of the spread of the quality itself. Spread, here and below, is the standard deviation.

draw
1975
correlation of the number with the quality, in the population drawn from
0.894

Fifty members, draw 1975. Each number is its quality plus an error of spread 0.50. In the population these fifty are drawn from, the number follows the quality with correlation 0.894, and it is never the quality itself.

2 That number ranks them, scored against the quality

Rank all fifty by the number, then score that order against the quality the page kept back. Two scores and three computations. Spearman’s rho is the correlation between the two lists of ranks; it is computed here from the ranks, and again from the formula in squared rank differences, and the two have to agree. That formula is exact only when nothing is tied, so on a tie the page refuses it. Kendall’s tau counts pairs: of the 1,225 pairs fifty members make, the number puts some in the order their quality would and some the other way round, and tau is the first count less the second, over all 1,225.

Spearman’s rho, from the ranks
0.895
the same, from the squared differences
0.895
Kendall’s tau
0.740
pairs in the right order
1,066 of 1,225

And what a population with this error gives on average, from closed forms rather than from this draw. For two normal variables with correlation r, Kendall’s tau averages (2/π) asin r. Spearman’s rho at fifty members averages Moran’s 1948 formula, 6/(π(n+1)) times (asin(r) + (n−2) asin(r/2)).

Kendall’s tau, on average
0.705
Spearman’s rho, on average
0.875

Ranked by the number, 1,066 of 1,225 pairs come out in the order their quality would put them, 87 in 100. A coin would get 50. Spearman's rho is 0.895 by ranks and 0.895 by squared differences, and Kendall's tau is 0.740. That order is the regularity Goodhart's sentence is about.

3 A cheaper way to raise the number, which skips the quality

Now a second way to raise the number, which does not go through the quality. It is cheaper than getting better, and it is open to some of them: by default the lower half by the number, who have the most reason to take it. Nobody lies. Whoever takes it really does score higher, by exactly the boost, and the number reports that faithfully. The lifted dots are drawn with the line they rose along.

could take it
25 of 50
took it
18
qualities that moved
0 of 50
numbers that are true readings
50 of 50

18 of the 25 who could took it, and each of their numbers rose by exactly 2.0. Not one quality moved. Every number is still a true reading of what happened: 50 of 50.

4 Rank them again, with every number still true

Rank again, and score again. Three columns: the ranking before, the ranking after the cheap route, and a control. The control is everybody getting better instead, with the number reading the improvement faithfully: the weakest gain most when the spread narrows, the strongest when it widens. It is scaled to end at exactly the spread the cheap route ends at, so a test that reads only the spread cannot tell the two apart.

measurebeforeafter the cheap routeeverybody improved instead
spread of the number1.1621.0001.000
Spearman's rho0.8950.4090.895
Kendall's tau0.7400.3010.740
pairs in order, of 1,2251,0667971,066
true readings, of 50505050

The spread narrowed by 14 per cent in both of the last two columns, to the same value. In one the order is untouched, 1,066 pairs right as before. In the other, 269 pairs that were in order are not any more, and 797 of 1,225 is 65 in 100, against a coin's 50. Nobody lied about a number.

The wording usually quoted is Marilyn Strathern’s, from 1997, and it is not Goodhart’s: “When a measure becomes a target, it ceases to be a good measure.” Her next sentence is the one this table measures: “The more a 2.1 examination performance becomes an expectation, the poorer it becomes as a discriminator of individual performances.”

PageRank is the same arrangement already on the roster: pages ranked by the links pointing at them, each weighted by the rank of the page it comes from: links anybody can count standing in for a quality nobody can see. Its page works through the formula, and nothing of it is re-derived here.

These ran in this browser when the page loaded. Each claim, whether it held, and the number behind it.

Each claim, whether it held, and the values behind it
claimheldmeasured
Spearman's rho comes out the same computed from the ranks and from the squared rank differencesyes0.895318 and 0.895318
Kendall's tau is the pairs in the right order less the pairs in the wrong one, over all 1225 pairsyes1066 right, 159 wrong, tau 0.740
over 400 populations the mean Kendall tau sits within 4 standard errors of the closed form, (2 / pi) asin ryesmean 0.7041, closed form 0.7048, 0.31 standard errors apart
and the mean Spearman rho sits within 4 standard errors of Moran's closed form for fiftyyesmean 0.8746, closed form 0.8749, 0.13 standard errors apart
after the cheap route every number is still a true reading: no quality moved, and each number moved by the boost or not at allyes50 of 50 hold; 18 took the route
and fewer pairs are in the right order than beforeyes1066 before, 797 after, of 1225
everybody improving instead ends at the same spread and leaves every pair where it wasyesspread 1.000 against 1.000; 1066 of 1225 in order, as before
a test that reads only the spread says the same about both, and only one of them kept its orderyesboth narrowed from 1.162 to 1.000
a route that every one of them takes changes no ranking at allyes50 of 50 took it; 1066 pairs in order, as before

What is real here, and what is not

1975 is when the papers were given; the volume is dated 1976

The Reserve Bank of Australia’s own bibliography lists Papers in Monetary Economics, volumes I and II, as published in Sydney in 1976, and describes it as a “Revised version of seven papers presented at the Conference in Monetary Economics, Sydney, July 1975, and two additional papers.” Beside Goodhart’s Problems of Monetary Management: The U.K. Experience it notes that “This paper contains the first reference to what became known as” Goodhart’s law. Chrystal and Mizen, and the 2021 editorial cited below, give the volume as 1975. The page carries July 1975, the date the papers were given, and both dates are firm. What the bibliography also says is that the printed papers were revised, so the printed wording is not certainly what was said in Sydney.

Which words, and where this page read them

The quotation in the lede is the form everybody repeats, and this page has not read it in any printing of Goodhart’s paper. There are three: the 1976 volume, a collection edited by Courakis in 1981, and Goodhart’s own Monetary Theory and Practice, Macmillan, 1984. Chrystal and Mizen name the second and third: the paper was published “in a volume edited by Courakis (1981) and then again in a volume of Charles’ own papers (Goodhart, 1984)”. None of the three could be read for this page, so the wording is checked against two secondary sources that agree on it. One is Chrystal and Mizen’s paper for Goodhart’s 2001 festschrift, which thanks him for comments; the other is a 2021 editorial in the Journal of Graduate Medical Education. In Chrystal and Mizen’s transcription the clause is introduced by the words “Ignoring Goodhart’s law”, and they write that “the statement of the law was a (jocular) aside rather than the main point of the paper”. Which printing those opening words come from is not established here.

A simulation with chosen parameters proves nothing about any institution

Fifty members, a quality drawn from a normal distribution, an error drawn from another, a boost of one fixed size, and a share who take it chosen by a uniform draw. Every one of those is a choice, and the page shows that a mechanism like this can break a ranking while every number stays true. It does not show that any real ranking has broken, or how often, or how fast. The page can score the order against the quality only because it invented the quality. In practice nobody has that column, which is why the false positive below matters.

A narrowing spread is a reason to look, and the control is why

The roamingpigs essay The Proof Was Never the Point puts it as advice: “A spread that bunches up is a reason to look, not a verdict”. The naive version of this test read a narrowing spread as the sign that a number had come loose from what it measured. Uniform improvement defeats it, so the page builds that case on purpose and makes it end at exactly the cheap route’s spread. The spread falls by the same amount in both columns, and only one of them keeps the order it had. The spread is not useless. It cannot, on its own, say which of the two happened.

The control keeps every pair in place because of how it is built

Everybody improving is drawn as one straight-line stretch applied to every quality and every number alike. When the spread has to narrow it is anchored at the best, so the weakest gain most; when it has to widen it is anchored at the worst. Either way nobody gets worse, and the value at the anchor does not move. A strictly increasing map moves nobody past anybody, and rank correlations ignore everything else, so Spearman’s rho and Kendall’s tau come out identical to before. That also shrinks each member’s measurement error with its score. If the improvement left the error at its old size, the same error would sit on a narrower spread and the order would erode a little, with no second route anywhere. That would be a third case the spread cannot separate from the other two, and this page does not draw it.

Only one of the ways a number fails

Manheim and Garrabrant sort the failures named after Goodhart into four kinds. Their first, regressional, is present as soon as stage 2 ranks by the number, with no second route at all: “When selecting for a proxy measure, you select not only for the true goal, but also for the difference between the proxy and the goal.” Their simple model of it is the one stage 1 draws, the goal plus normal noise. The second route in stage 3 is a different failure: the number is raised by a path that does not pass through the quality. This page draws that one and does not claim to cover the rest.

Not the first statement of the idea, and the page does not say it is

Chrystal and Mizen weigh Goodhart’s law against the Lucas critique and conclude that if the two are the same thing, Lucas almost certainly said it first. Manheim and Garrabrant call Campbell’s law the formulation “which arguably has scholarly precedence”. Goodhart’s is the one with his name on it and a firm date, which is all this page needs.

The closed forms are checked here, not trusted

Kendall’s tau for a pair of normal variables with correlation r averages (2/π) asin r, and Spearman’s rho at sample size n averages Moran’s 1948 formula. Both are set out in a tutorial by de Winter, Gosling and Potter, and the page holds each against a sweep it runs when it loads: four hundred populations from fixed seeds, with the mean required to sit within four standard errors of the formula. The tutorial’s own worked values for a correlation of .8, .688 at five members and .758 at twenty, come out of the same formula here.

Sound: no

Asked and answered, so it is not reopened. Nothing measured here has a duration. A tone per pair out of order would be a decoration of a count already printed.

Sources