Goodhart
In July 1975 Charles Goodhart gave two papers at a Reserve Bank of Australia conference in Sydney. One of them carries, as an aside, the sentence his name is now attached to: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes”. This page is one reading of that sentence made into something you can operate. Fifty members have a quality nobody can see and a number read off it, and the number ranks them well. Then a cheaper way to raise the number arrives, some of them take it, and the ranking breaks while every number in it stays true. The easy mistake is in the test. A spread that narrows proves nothing on its own, because everybody getting better can narrow it just as far, and this page shows both.
New to a number standing in for a quality? Start here
Most of what gets ranked cannot be measured directly. How good a student is, how useful a page is, how sound a bank is: nobody can read those off a dial, so people agree on a number that tends to go with them, an exam mark, a count of links, a ratio. The number is a proxy, and it keeps its order only as long as the way to raise it is to have more of the thing it stands for.
To ask whether a proxy still works you do not need its values to be right, only its order. Take every pair of members and ask whether the number puts them the same way round as the quality would. Count the pairs it puts the right way round and the pairs it puts the wrong way round, divide the difference by all the pairs, and you have Kendall's tau; correlate the two lists of ranks and you have Spearman's rho. Neither cares how big the numbers are, which is why a whole population can improve without either moving.
This page holds the quality as well as the number, because it made both up, and that is the one thing no real ranking can do.
Machines here that come first: Sampling.
A ranking, a second route to the number, and the control that must pass
1 A quality nobody can see, and a number read off it
Fifty members, and each has a quality. Nobody outside this page can see it; the page can only because it made it up. What everybody can see is a number, and the number is the quality plus an error of measurement. The error is the knob, measured in units of the spread of the quality itself. Spread, here and below, is the standard deviation.
- draw
- 1975
- correlation of the number with the quality, in the population drawn from
- 0.894
Fifty members, draw 1975. Each number is its quality plus an error of spread 0.50. In the population these fifty are drawn from, the number follows the quality with correlation 0.894, and it is never the quality itself.
2 That number ranks them, scored against the quality
Rank all fifty by the number, then score that order against the quality the page kept back. Two scores and three computations. Spearman’s rho is the correlation between the two lists of ranks; it is computed here from the ranks, and again from the formula in squared rank differences, and the two have to agree. That formula is exact only when nothing is tied, so on a tie the page refuses it. Kendall’s tau counts pairs: of the 1,225 pairs fifty members make, the number puts some in the order their quality would and some the other way round, and tau is the first count less the second, over all 1,225.
- Spearman’s rho, from the ranks
- 0.895
- the same, from the squared differences
- 0.895
- Kendall’s tau
- 0.740
- pairs in the right order
- 1,066 of 1,225
And what a population with this error gives on average, from closed forms rather than from this draw. For two normal variables with correlation r, Kendall’s tau averages (2/π) asin r. Spearman’s rho at fifty members averages Moran’s 1948 formula, 6/(π(n+1)) times (asin(r) + (n−2) asin(r/2)).
- Kendall’s tau, on average
- 0.705
- Spearman’s rho, on average
- 0.875
Ranked by the number, 1,066 of 1,225 pairs come out in the order their quality would put them, 87 in 100. A coin would get 50. Spearman's rho is 0.895 by ranks and 0.895 by squared differences, and Kendall's tau is 0.740. That order is the regularity Goodhart's sentence is about.
3 A cheaper way to raise the number, which skips the quality
Now a second way to raise the number, which does not go through the quality. It is cheaper than getting better, and it is open to some of them: by default the lower half by the number, who have the most reason to take it. Nobody lies. Whoever takes it really does score higher, by exactly the boost, and the number reports that faithfully. The lifted dots are drawn with the line they rose along.
- could take it
- 25 of 50
- took it
- 18
- qualities that moved
- 0 of 50
- numbers that are true readings
- 50 of 50
18 of the 25 who could took it, and each of their numbers rose by exactly 2.0. Not one quality moved. Every number is still a true reading of what happened: 50 of 50.
4 Rank them again, with every number still true
Rank again, and score again. Three columns: the ranking before, the ranking after the cheap route, and a control. The control is everybody getting better instead, with the number reading the improvement faithfully: the weakest gain most when the spread narrows, the strongest when it widens. It is scaled to end at exactly the spread the cheap route ends at, so a test that reads only the spread cannot tell the two apart.
| measure | before | after the cheap route | everybody improved instead |
|---|---|---|---|
| spread of the number | 1.162 | 1.000 | 1.000 |
| Spearman's rho | 0.895 | 0.409 | 0.895 |
| Kendall's tau | 0.740 | 0.301 | 0.740 |
| pairs in order, of 1,225 | 1,066 | 797 | 1,066 |
| true readings, of 50 | 50 | 50 | 50 |
The spread narrowed by 14 per cent in both of the last two columns, to the same value. In one the order is untouched, 1,066 pairs right as before. In the other, 269 pairs that were in order are not any more, and 797 of 1,225 is 65 in 100, against a coin's 50. Nobody lied about a number.
The wording usually quoted is Marilyn Strathern’s, from 1997, and it is not Goodhart’s: “When a measure becomes a target, it ceases to be a good measure.” Her next sentence is the one this table measures: “The more a 2.1 examination performance becomes an expectation, the poorer it becomes as a discriminator of individual performances.”
PageRank is the same arrangement already on the roster: pages ranked by the links pointing at them, each weighted by the rank of the page it comes from: links anybody can count standing in for a quality nobody can see. Its page works through the formula, and nothing of it is re-derived here.
These ran in this browser when the page loaded. Each claim, whether it held, and the number behind it.
| claim | held | measured |
|---|---|---|
| Spearman's rho comes out the same computed from the ranks and from the squared rank differences | yes | 0.895318 and 0.895318 |
| Kendall's tau is the pairs in the right order less the pairs in the wrong one, over all 1225 pairs | yes | 1066 right, 159 wrong, tau 0.740 |
| over 400 populations the mean Kendall tau sits within 4 standard errors of the closed form, (2 / pi) asin r | yes | mean 0.7041, closed form 0.7048, 0.31 standard errors apart |
| and the mean Spearman rho sits within 4 standard errors of Moran's closed form for fifty | yes | mean 0.8746, closed form 0.8749, 0.13 standard errors apart |
| after the cheap route every number is still a true reading: no quality moved, and each number moved by the boost or not at all | yes | 50 of 50 hold; 18 took the route |
| and fewer pairs are in the right order than before | yes | 1066 before, 797 after, of 1225 |
| everybody improving instead ends at the same spread and leaves every pair where it was | yes | spread 1.000 against 1.000; 1066 of 1225 in order, as before |
| a test that reads only the spread says the same about both, and only one of them kept its order | yes | both narrowed from 1.162 to 1.000 |
| a route that every one of them takes changes no ranking at all | yes | 50 of 50 took it; 1066 pairs in order, as before |
What is real here, and what is not
1975 is when the papers were given; the volume is dated 1976
The Reserve Bank of Australia’s own bibliography lists Papers in Monetary Economics, volumes I and II, as published in Sydney in 1976, and describes it as a “Revised version of seven papers presented at the Conference in Monetary Economics, Sydney, July 1975, and two additional papers.” Beside Goodhart’s Problems of Monetary Management: The U.K. Experience it notes that “This paper contains the first reference to what became known as” Goodhart’s law. Chrystal and Mizen, and the 2021 editorial cited below, give the volume as 1975. The page carries July 1975, the date the papers were given, and both dates are firm. What the bibliography also says is that the printed papers were revised, so the printed wording is not certainly what was said in Sydney.
Which words, and where this page read them
The quotation in the lede is the form everybody repeats, and this page has not read it in any printing of Goodhart’s paper. There are three: the 1976 volume, a collection edited by Courakis in 1981, and Goodhart’s own Monetary Theory and Practice, Macmillan, 1984. Chrystal and Mizen name the second and third: the paper was published “in a volume edited by Courakis (1981) and then again in a volume of Charles’ own papers (Goodhart, 1984)”. None of the three could be read for this page, so the wording is checked against two secondary sources that agree on it. One is Chrystal and Mizen’s paper for Goodhart’s 2001 festschrift, which thanks him for comments; the other is a 2021 editorial in the Journal of Graduate Medical Education. In Chrystal and Mizen’s transcription the clause is introduced by the words “Ignoring Goodhart’s law”, and they write that “the statement of the law was a (jocular) aside rather than the main point of the paper”. Which printing those opening words come from is not established here.
A simulation with chosen parameters proves nothing about any institution
Fifty members, a quality drawn from a normal distribution, an error drawn from another, a boost of one fixed size, and a share who take it chosen by a uniform draw. Every one of those is a choice, and the page shows that a mechanism like this can break a ranking while every number stays true. It does not show that any real ranking has broken, or how often, or how fast. The page can score the order against the quality only because it invented the quality. In practice nobody has that column, which is why the false positive below matters.
A narrowing spread is a reason to look, and the control is why
The roamingpigs essay The Proof Was Never the Point puts it as advice: “A spread that bunches up is a reason to look, not a verdict”. The naive version of this test read a narrowing spread as the sign that a number had come loose from what it measured. Uniform improvement defeats it, so the page builds that case on purpose and makes it end at exactly the cheap route’s spread. The spread falls by the same amount in both columns, and only one of them keeps the order it had. The spread is not useless. It cannot, on its own, say which of the two happened.
The control keeps every pair in place because of how it is built
Everybody improving is drawn as one straight-line stretch applied to every quality and every number alike. When the spread has to narrow it is anchored at the best, so the weakest gain most; when it has to widen it is anchored at the worst. Either way nobody gets worse, and the value at the anchor does not move. A strictly increasing map moves nobody past anybody, and rank correlations ignore everything else, so Spearman’s rho and Kendall’s tau come out identical to before. That also shrinks each member’s measurement error with its score. If the improvement left the error at its old size, the same error would sit on a narrower spread and the order would erode a little, with no second route anywhere. That would be a third case the spread cannot separate from the other two, and this page does not draw it.
Only one of the ways a number fails
Manheim and Garrabrant sort the failures named after Goodhart into four kinds. Their first, regressional, is present as soon as stage 2 ranks by the number, with no second route at all: “When selecting for a proxy measure, you select not only for the true goal, but also for the difference between the proxy and the goal.” Their simple model of it is the one stage 1 draws, the goal plus normal noise. The second route in stage 3 is a different failure: the number is raised by a path that does not pass through the quality. This page draws that one and does not claim to cover the rest.
Not the first statement of the idea, and the page does not say it is
Chrystal and Mizen weigh Goodhart’s law against the Lucas critique and conclude that if the two are the same thing, Lucas almost certainly said it first. Manheim and Garrabrant call Campbell’s law the formulation “which arguably has scholarly precedence”. Goodhart’s is the one with his name on it and a firm date, which is all this page needs.
The closed forms are checked here, not trusted
Kendall’s tau for a pair of normal variables with correlation r averages (2/π) asin r, and Spearman’s rho at sample size n averages Moran’s 1948 formula. Both are set out in a tutorial by de Winter, Gosling and Potter, and the page holds each against a sweep it runs when it loads: four hundred populations from fixed seeds, with the mean required to sit within four standard errors of the formula. The tutorial’s own worked values for a correlation of .8, .688 at five members and .758 at twenty, come out of the same formula here.
Sound: no
Asked and answered, so it is not reopened. Nothing measured here has a duration. A tone per pair out of order would be a decoration of a count already printed.
Sources
- Suzanna Chiang and Michael Power, A Select Bibliography of Published Research by the Reserve Bank of Australia: 1969–1990, Research Discussion Paper 9013, December 1990, conference volumes. The July 1975 conference, the 1976 imprint, and the note that Goodhart’s paper is the first reference to his law.
- K. Alec Chrystal and Paul D. Mizen, Goodhart’s Law: Its Origins, Meaning and Implications for Monetary Policy, dated 12 November 2001, prepared for the festschrift in Goodhart’s honour at the Bank of England; read in a third party’s online copy. The clause in its sentence, the 1981 and 1984 reprints, and the Lucas critique.
- Christopher Mattson, Reamer L. Bushardt and Anthony R. Artino, Jr., When a Measure Becomes a Target, It Ceases to be a Good Measure, Journal of Graduate Medical Education 13, 2021. The second witness to the wording of Goodhart’s clause.
- Marilyn Strathern, ‘Improving ratings’: audit in the British University system, European Review 5, 305–321, 1997, page 308. The rewording everybody quotes, and the sentence after it.
- J. C. F. de Winter, S. D. Gosling and J. Potter, Comparing the Pearson and Spearman Correlation Coefficients Across Distributions and Sample Sizes, Psychological Methods 21, 273–290, 2016; read in the authors’ arXiv copy of 2024. The closed forms for Kendall’s tau and for Spearman’s rho at a given sample size, with worked values.
- David Manheim and Scott Garrabrant, Categorizing Variants of Goodhart’s Law, 2018, revised 2019. Four ways a proxy fails, of which this page draws one.
- The Proof Was Never the Point, roamingpigs.com. Where the naive spread test failed on uniform improvement, which is why the control here is required.
- Logical Art, the studio this belongs to.