I spent a good part of a year in graduate school arguing with another doctoral student about whether basketball players get hot.
His position was that the question was closed. Mine was that it could not be closed, because if it were closed somebody would have made money on it by now. We were both partly right, we were both arguing about the wrong thing, and it took a paper published thirty-three years after the study we were arguing about to show why.
Here is where this lands, so nobody has to wait for it. The study we were arguing about measured a psychological bias using a tool that was itself statistically biased. A statistical bias and a psychological bias have close to nothing in common, and telling those two apart is what this story is for. Getting there needs a coin and about five minutes.
What the famous study found
In 1985, three psychologists asked a hundred basketball fans whether a player has a better chance of making a shot right after he has made his last two or three. Ninety-one said yes, and put numbers on it: a 50 percent shooter would hit about 61 percent after a make and 42 percent after a miss.
Then the researchers went and looked. Philadelphia 76ers field goals for a season, Boston Celtics free throws for two. A make did not predict a make.
There is an obvious objection and they had thought of it. In a real game a player who just hit two starts taking harder shots and the defense starts crowding him, and either would hide a hot hand that was really there. So they ran a controlled experiment. Twenty-six Cornell players, a hundred shots each, from a distance set individually so every player would hit about half. No defenders, no choosing your spot.
Same answer. And before each shot the players bet on themselves, and the bets tracked what they had just done rather than what they were about to do.
That is where the phrase "hot hand fallacy" entered the language.
My side of it
I never disputed the shooting data. What bothered me was everybody else.
An NBA head coach is paid several million dollars a year, in large part for seeing patterns in a game and acting on them before the other guy does. There are thirty of them and they have every frame of film ever shot. All of them believed in the hot hand, and so did the players. If the belief were an illusion and nothing else, a standing edge was available to any franchise willing to stop believing it, and it had been sitting on the table since 1985.
His answer was good and I did not have much for it. People believe wrong things confidently, in groups, for decades, and being well paid does not fix that. Coaches believing something is a fact about coaches.
I owe him more than I gave him at the time, because that argument has since been tested and it mostly comes out his way. The cleanest finding against me is that after a single make, a player is more likely to take his team's next shot, takes it from farther out, and is less likely to make it. Acting on the belief costs points whether or not the belief is true.
A second study found that defenders do crowd a hot shooter and that the crowding is not vindicated, since the supposedly hot players do not shoot better than usual. That settles less than it looks like: a shooter who draws tighter defense and holds his percentage is doing what a real hot hand under pressure would look like. A later study using tracking data put the hot hand at one to two and a half points once shot difficulty was accounted for, which is small and is not nothing.
So I lost the argument I was making. Something was still wrong with the 1985 result, and it had nothing to do with coaches.
Flip a coin three times
Flip a coin three times. Every time you get a head, write down whatever comes up on the very next flip. Then work out what fraction of the flips you wrote down were heads. You would expect that to settle at one half. It settles at about 0.42.
There are only eight ways three flips can come out, so you can check this without flipping anything. In each sequence, find every flip that comes up heads, record whatever comes immediately after it, and count how many of the recorded flips were themselves heads.
H H H. The first flip is heads, so you record the second, which is heads. The second flip is heads, so you record the third, which is heads. You recorded two flips and both were heads. Two out of two, which is 1.
H H T. The first is heads, so you record the second, which is heads. The second is heads, so you record the third, which is tails. You recorded two flips and one of them was heads. One out of two, which is one half.
H T H. Only the first flip is heads, so you record only the second, which is tails. You recorded one flip and it was not heads. Zero out of one, which is 0.
H T T. Only the first is heads, so you record only the second, which is tails. One flip recorded, zero heads. That is 0.
T H H. Only the second flip is heads, so you record only the third, which is heads. One flip recorded, one head. One out of one, which is 1.
T H T. Only the second is heads, so you record only the third, which is tails. One flip recorded, zero heads. That is 0.
T T H. The only heads is the third flip, and there is no fourth flip to record. Nothing gets recorded, so this sequence gives you no number at all.
T T T. No heads anywhere, so nothing gets recorded. No number.
Six of the eight sequences gave you a number. Those six numbers, in the order above, are 1, one half, 0, 0, 1, and 0. Adding them together gives two and a half. Two and a half divided by six is five twelfths, which is 0.4167.
Now do it the other way. Ignore the sequences and count every flip you recorded, all eight of them across all eight sequences. Four are heads. Four out of eight is one half exactly, which is what the coin owes you.
Same coin, same procedure, two answers. Averaging the sequences gives 0.4167. Pooling the flips gives 0.5. The gap between those two numbers is the entire subject of the rest of this post.
Why the averaging comes out low
Inside a finite record, the heads you already used cannot land again. Every time you condition on a head you spend one, and the ones left to fill the spot you are watching are fewer than the sequence's own rate.
Pooling escapes it by weighting each recorded flip equally instead of weighting each sequence equally, and any single recorded flip has a fair one-in-two chance no matter which sequence it sits in.
That is the whole mechanism. It is worth doing on a napkin, because reading the eight sequences is not the same as watching five twelfths come out of your own handwriting.
Two economists, Joshua Miller and Adam Sanjurjo, proved this, worked out how large the effect is, and pointed it at the 1985 study.
What that does to the finding
The 1985 measure was the coin procedure wearing a jersey. What gets counted is one number per shooter: his shooting percentage on the shots that followed three makes, minus his percentage on the shots that followed three misses. Then those twenty-six numbers get averaged together. A proportion worked out inside each record, then averaged across records, which is the coin exactly.
Run the same arithmetic on a shooter. A man who makes fifty of a hundred has, conditional on a shot that followed a make, about forty-nine makes left among ninety-nine remaining shots. That shortfall is tiny at one shot. Condition on three makes in a row rather than one and it grows, which is the part Miller and Sanjurjo proved and put a number on.
Under the 1985 design a player with no hot hand whatsoever would come out about eight percentage points on the cold side of that measure. The instrument read low before anyone picked up a ball, so a player had to be genuinely hot to score as merely even.
The Cornell shooters averaged about three percentage points on the hot side, small enough to read as noise. Add back what the measure had subtracted and that same average is about thirteen points on the hot side. Which is a statement about the measure, not a claim that anybody was thirteen points hot.
The sentence with both meanings in it
A study measuring a psychological bias used a statistically biased instrument.
A biased instrument and a biased judgment share almost nothing, and the two words sit next to each other constantly.
A statistically biased instrument reads systematically off. A bathroom scale that says two pounds heavy is biased. The scale has not made a mistake and holds no opinions, and you fix it with arithmetic. A biased estimator is sometimes the one you want, since being off in a known direction can beat being scattered.
A psychologically biased judgment departs from a rule about how the judging should have gone. That sense carries a whiff of error and often a whiff of character. It is the sense the 1985 paper meant when it called belief in the hot hand "a powerful and widely shared cognitive illusion."
What the two share is one narrow property. Systematic rather than random, off in a direction rather than scattered. Everything else differs, starting with whether anybody did anything wrong. Statistical bias is a fact about an instrument. Psychological bias is a verdict on a person. The two get set down in adjacent sentences all day long.
What remains after the correction
The 1985 paper made two claims. The shooting data showed no hot hand, and fans wildly over-perceived one. Miller and Sanjurjo challenge the first claim and leave the second standing. Free throws come in pairs, so the arithmetic problem never arises there. The eight-point gap fans expect is still close to nothing.
The psychological finding survives the correction to the instrument. A reader who leaves this post believing the hot hand fallacy was overturned has been misled. The sentence I built the post around, that a study measuring a psychological bias used a statistically biased instrument, can mislead in the same way. The tracking estimate of one to two and a half points also means that coaches’ belief is not purely an illusion. I had underestimated that. What I got wrong was the size of the edge available to anyone who stopped believing it.
Where it stands
Not where either of us would have guessed.
Miller and Sanjurjo do not claim that players get hot in games. They argue that in-game data cannot settle it, because the defense adjusts, which is the same objection the 1985 authors raised themselves and the reason they built the Cornell experiment in the first place. Their claim is narrower than the headlines. The data that had been read as proving the belief was a fallacy does not prove that.
Others think even the narrow claim outruns the evidence. One reanalysis points out that this whole dispute runs on twenty-six people taking a hundred shots each, and that once you correct for testing twenty-six shooters at once, exactly one of them shoots in a way you can tell apart from a coin.
And the over-perception may survive all of it. Free throws come in pairs, so there is no long record to select streaks out of and the arithmetic problem above does not arise. Fans still expect a gap of about eight points between a shooter who just made one and a shooter who just missed. The real gap is close to nothing.
So the argument I had in graduate school is still open. What changed was the terms of it.