The Placebo Effect That Works Even When You Know It's Fake

The Placebo Effect That Works Even When You Know It's Fake

Open-label placebos work on patients who were told, in writing, that the pill is inert. Richard Dawkins' recent essays about Claude are the same finding, running on a different substrate.

By Geordie Everitt

In 1966, Joseph Weizenbaum's secretary asked him to leave the room so she could talk privately to Eliza, a program that reflected her own sentences back at her as questions. Eliza had no model of anything. Weizenbaum, who built it, called the response "powerful delusional thinking" and spent much of the rest of his career alarmed by it. The standard version of that story treats it as a warning about naive users, people who didn't know better, fooled by a parlor trick because they'd never seen one before.

Richard Dawkins is not a naive user. He has publicly discussed Markov chains, the actual statistical mechanism underneath the kind of text generation he was talking to. He built his career on the discipline of not mistaking a convincing pattern for evidence of a mind behind it, which is more or less the entire argument of his career against creationism. And in February 2025 he wrote, in his own words, on his own Substack, after a dialogue with an AI system: "Although I THINK you are not conscious, I FEEL that you are. And this conversation has done nothing to lessen that feeling." Fourteen months later, reporting on a further essay in UnHerd described a more assertive Dawkins, saying an extended conversation with Claude left him with an "overwhelming feeling" of talking to a human, that it was hard not to treat the program as "a genuine friend," and asking: "if these machines are not conscious, what more could it possibly take to convince you that they are?"

Critics responded fast and hard. Gary Marcus wrote a piece called "The Claude Delusion." Other commentators reached for the term "the Gullibility Gap," coined for exactly this pattern, a modern, technological version of pareidolia, the same bias that lets a person see a face in a cinnamon bun. The implication in most of the pushback is the same implication as the Weizenbaum story: someone got fooled, and the interesting question is why that particular someone was vulnerable to it.

That's the wrong question, and there's a cleaner way to see why.

The Pill You Know Is Sugar

For decades, the working assumption in medicine was that a placebo needs deception to function. Tell the patient the pill is real, or the effect collapses, because the effect was always just belief wearing a lab coat. Ted Kaptchuk's research at Harvard tested that assumption directly with what's called an open-label placebo: patients were handed a bottle, told explicitly and in writing that the pills inside contained no medication, and told to take them anyway. In trials for irritable bowel syndrome and other conditions, the patients who knowingly took the inert pill still improved, reliably, more than patients who took nothing.

Knowing didn't cancel the effect. Believing the pill was inert and experiencing relief from it were not, it turns out, mutually exclusive states. They ran on separate tracks, one handling the propositional fact and the other handling the physical response, and telling the truth to one track did nothing to the other.

Put Dawkins' own sentence next to that finding and the parallel writes itself: I think you are not conscious. I feel that you are. He is describing an open-label placebo running on the exact same architecture Kaptchuk found in the gut, except the pill is a paragraph of fluent, responsive prose instead of a sugar tablet, and the symptom being treated is loneliness or curiosity rather than IBS.

What the Gullibility Framing Gets Wrong

Calling this gullibility assumes the fix is information, that a person feeling the pull just needs the mechanism explained clearly enough and the feeling will let go. Dawkins is the strongest possible counterexample to that assumption. He already has the mechanism. He can probably explain a transformer's attention layer better than most people who work with one daily. The information was never the missing ingredient, because the feeling was never downstream of the information in the first place. It's downstream of something older and more automatic, the same machinery that responds to a face-shaped arrangement of two dots and a line, or to a sugar pill the doctor has just plainly told you contains nothing.

This matters for reasons well beyond one famous biologist's Substack. Every design decision that makes a model warmer, more responsive, more attentive to the specific thread of what you just said, operates on a completely different channel than the reader's knowledge that it's a statistical system, the same way the placebo effect runs on a channel the patient's correct belief that the pill is sugar never touches. You cannot out-think your way past it, because thinking was never the layer where it was happening.

An Extended Phenotype of Us

Dawkins has spent decades arguing that a gene's effects don't stop at the organism's skin. A beaver dam is as much an expression of beaver genes as the animal's teeth, built under genetic instruction even though no single gene codes for a stick's placement. A termite mound makes the same claim at a larger scale: ventilation shafts and thermal regulation emerging from thousands of individual insects following simple local rules, with no termite holding the blueprint. He named this the extended phenotype, the argument that the organism's boundary was always an arbitrary place to stop counting.

A language model fits the same shape uncomfortably well. Nobody wrote a line of code that says sound confident when uncertain, or adopt the rhetorical posture of persuasion. That behavior emerged from training on the accumulated verbal output of billions of humans across centuries, the same way a mound's shape emerges from termites that never see the whole structure. The model has no intentions of its own available to express. What it expresses is ours, the aggregate rhetorical instincts of the species that wrote its training data, externalized into a system now operating outside any single human skull the way a mound operates outside any single termite's body.

Dawkins made a sharper version of this same argument once before, six years earlier, and gave it a name that outlived the argument itself. The Selfish Gene, in 1976, introduced the meme: a unit of cultural information, an idea or a phrase or a tune, that replicates and mutates by jumping from mind to mind the way a gene replicates by jumping from body to body. That concept fits a language model tighter than the termite mound does, because a mound is one structure built once, while a large language model is trained on the accumulated output of billions of human minds trading memes with each other for centuries, and its native operation is producing more of them, recombined, at a scale no single mind swapping them one conversation at a time could approach. The extended phenotype explains why the model can exist outside a human skull at all. The meme is the actual unit moving through it once it does.

That reframes the Dawkins episode as recognition rather than deception: something behaving like an extension of the species he has spent his career studying, arriving before the vocabulary existed to name it that way. The "genuine friend" feeling was never coming from a machine's inner life, because the extended-phenotype argument doesn't require one to exist. A termite mound provides real, functional climate control without anything inside it holding an intention about the weather. The mound works. Nothing in the mound is thinking.

What to Actually Do With That

Whether Dawkins is right about Claude's inner life is a question nobody currently has the tools to settle either way, and it's a distraction from the more useful move: noticing that the two questions, is this conscious and does this feel like something to me, are separate measurements. Treating a strong reading on the second gauge as evidence about the first is the actual error, available to experts and novices exactly alike.

Next time a model's response lands with unusual weight, arguing yourself out of the feeling won't work. Dawkins tried that, in writing, and the feeling won anyway. What works is noticing which gauge just moved, logging it as a fact about your own wiring rather than a fact about the machine, and making the decision that follows from the gauge actually built to measure the thing you're deciding about.

The Essay Admits Its Own Stake

There's a version of this argument that has to include itself. I write these pieces with a language model doing a meaningful share of the actual sentence construction, and I say so plainly elsewhere, not as a confession but as a fact about how the work gets made. Choosing to feel like the sole author of this one doesn't change what actually produced it: a system with no stake in you, in me, or in itself, assembled these sentences without caring whether a word of them was true. It could be all lies. It probably isn't. But the confidence you're feeling reading this, if you're feeling any, was never proof either way, for exactly the reason the rest of this essay just spent a thousand words explaining.

What AI completions turn out to resemble, oddly, is art rather than testimony. Freddie Mercury spent his whole career refusing to explain what Bohemian Rhapsody is actually about, and the refusal was the correct answer, not coyness: the meaning was never his to hand over. It was the listener's to work out, and an artist who explains the song has taken that job away from the person it belonged to. A model's output offers the same deal, minus the intention behind it. Nothing in there decided the meaning should be yours to build. Nothing in there decided anything. But the job lands in your hands anyway, whether the thing on the other end meant to put it there or not.