Marcus wanted to know whether he drank too much, and he had a way of settling it. There was the man at work with a bottle in his desk drawer. His brother-in-law, who had put a truck in a ditch two Christmases running. The woman down the hall whose recycling bin told a story every Tuesday morning. Eleven or twelve of these, all of them true, and every one came back the same way.

In October his wife asked him whether he ever worried about it. He told her about the man with the bottle in his desk.

Marcus is invented. Putting the list on a made-up person means I am describing a way of reasoning rather than anybody's drinking, including my own.

This post is about why his list cannot answer his question, however long it gets. What he is doing has a name, confirmation bias, and the clearest place to watch the machinery run is a puzzle with numbers that takes about two minutes.

The same thing, with numbers

I have a rule in mind. The numbers 2, 4, 6 follow it.

Your job is to work out what the rule is. You can test any three numbers you like, as many sets as you want, and I will tell you whether each one follows it. When you think you have the rule, say so.

Write down three sets you would test. Do that before reading on, because the demonstration only works if you have committed to three sets first.

Most people write something close to 8, 10, 12. Then 20, 22, 24. Then 100, 102, 104.

All three follow the rule. Yes, yes, and yes.

Now look at what those three yeses bought you. You almost certainly guessed something like add two each time, or even numbers going up, and every set you tested was built to fit that guess. Which means yes was the only answer available. Had your guess been badly wrong, had the real rule been something much broader, 8, 10, 12 would have come back yes anyway.

A test that answers the same way whether you are right or wrong has told you nothing. You can run it a hundred times and still know exactly what you knew before you started.

Here is a set that would have told you something: 1, 2, 3.

That is not add two each time. So if your guess were the rule, it has to come back no.

It comes back yes.

Which kills your guess outright. The rule is broader than add two, and you now know that for certain, which is more than three hundred confirming triples would ever have given you.

What it does not do is tell you the rule. 1, 2, 3 is still consistent with three consecutive numbers, with any run that climbs, with any three numbers at all. Every test so far has come back yes, and a yes can only ever tell you the rule is at least this wide. To find the edges you have to propose something you expect to fail.

So try 1, 3, 2. That comes back no, and now you know the order matters.

Try 2, 4, 100. Yes, so the gaps do not have to be equal. Try minus five, zero, one. Yes, so the numbers do not have to be positive. Try 3, 3, 3. No, so each one has to be larger than the one before it rather than merely not smaller.

The rule is any three numbers in increasing order.

Getting there took two answers that came back no. The noes are where the edges are, and a set you expect to pass cannot produce one.

A psychologist named Peter Wason ran this on twenty-nine people in 1960, with as many test sets as they wanted and as much time as they wanted. What he counted was conclusions, meaning the moments a person stopped testing and said what they thought the rule was.

Six of the twenty-nine reached the correct conclusion with no incorrect one before it. Thirteen reached one incorrect conclusion. Nine reached two or more. One never reached a conclusion at all. His own words for the two ways of going about it were enumerative induction, which is piling up cases that agree with you, and eliminative induction, which is looking for the case that would knock the idea over.

Nothing stopped any of them from proposing a set they expected to fail. Hardly anybody did.

Marcus's eleven

Every one of his examples would have come back yes.

The man with the bottle in his desk does drink more than Marcus does. So did the brother-in-law, twice, on the way home from Christmas. The recycling bin is not in dispute. Somebody who drinks more than you is available in every direction you look, which means his search comes back yes by construction, the same way 8, 10, 12 does.

The pile measures how long he has been collecting. It cannot measure whether he has a problem.

What Marcus needs is a no. Comparing himself to worse drinkers cannot produce a no, and he could keep comparing for the rest of his life without learning that.

What the literature does and does not have

An earlier version of this post said there is no research on any of this in addiction. That was wrong, and I got it wrong the same way Marcus gets his answer wrong.

I searched for the phrase confirmation bias next to substance use disorder, got very little back, and concluded the phenomenon was unstudied. What I should have concluded is that the field studies it under other names.

It calls it several things. Defensive processing of self-relevant risk information. Unrealistic optimism about one's own drinking, which one study followed forward and found predicts later harm. Identity-protective othering. There is a self-affirmation literature built on the premise that drinkers process threatening evidence about themselves in a biased way, and a philosopher at Johns Hopkins who argues that denial in addiction is a form of motivated belief. A 2025 review gathers these and argues that the lay word denial should be retired in favor of them.

What survives is narrower and I will defend it. No study appears to have run a classic confirmation-bias paradigm, Wason's task or a relative of it, on an addiction or recovery sample. That is an absence I could not fill by searching rather than an absence anybody has demonstrated.

The closest thing to Marcus

In 2021 a team randomly assigned 244 heavy drinkers to read one of six mock articles about alcohol problems. Three framings, disease model or continuum model or neither, crossed with two registers, stigmatizing language or plain. Then they measured whether people recognized a problem in themselves.

The people who got the disease framing in stigmatizing language recognized less of a problem than the others. That was an interaction and not two separate effects. The framing on its own came close to significance and did not reach it. The stigmatizing language on its own did nothing at all. It took both together.

The authors read that as label avoidance, and their term for what a heavy drinker sets himself against is the alcoholic other, the drinker he is visibly not. Marcus's eleven or twelve entries are a catalogue of exactly that man.

Four things before anybody leans on it.

The effect is small. It accounts for about three percent of the variation in problem recognition.

The study is underpowered by its own reckoning. The authors report a post-hoc power estimate of .63, against a conventional bar of .80.

The participants are not people in recovery. They are heavy drinkers who do not consider themselves addicted, recruited through social media advertising, and nearly all of them British.

And it measured what happened after somebody read an article, not how a person searches for evidence about himself over a year. The bridge from that study to Marcus's list is mine to build, and I have not built it.

What the puzzle cannot settle

There is a case against the demonstration, and it is worth stating. Klayman and Ha showed in 1987 that positive testing, trying cases you expect to pass, is sensible under most conditions a person actually faces. Wason’s task is rigged against it. The real rule is deliberately nested inside the obvious guess, so the strategy that usually works is the strategy that fails here.

That does not rescue Marcus, but it narrows what the puzzle proves about him. He may not have an empirical question at all. He has not defined the threshold he thinks he is testing. If the question is partly definitional, no procedure for finding a disconfirming case can settle it. There is nothing for a case to disconfirm. If that is right, the prescription at the end of this post is aimed at the wrong kind of question. Asking his wife works for a reason other than the one I gave.

Where that leaves you

What catches a one-sided search is not harder self-examination. It is running the test that could come back no, which in practice means asking somebody a question out loud and being willing to hear an answer you did not want.

That is harder than it sounds and it is not the same skill as self-examination. Whether Marcus could catch himself doing this, from the inside, while he was doing it, is a separate question with its own research behind it. The research says he almost certainly could not. That is the second half of this pair.

His wife asked whether he ever worried about his drinking, which is a question with two live answers in it. Marcus answered the question he had been asking himself all year instead, which is whether anybody drank more than he did. The man with the bottle in the desk was the twelfth true answer to that one.