The AI Lie: Why It Will Never Change

The AI Lie: Why It Will Never Change

AI is a remorseless liar, trained on a mountain of text whose whole purpose was persuasion, not accuracy. No amount of prompting fixes that on its own. You have to stop enabling it.

By Geordie Everitt

Ask an image model to draw an analog clock. Any time you like. It will draw ten past ten, or twenty to two, and it will keep doing this no matter how you word the request, because those two hand positions are what watch advertising has used for a century to frame the brand name and look symmetrical doing it. The model has seen millions of clocks in training and an overwhelming share of them say the same thing. Ask for three fifteen and you get ten ten with a different label. Sometimes there is no amount of prompting that undoes it. The groove is worn too deep for the instruction to reach.

That failure is trivial. Nobody's hurt by a wrong clock face. It's useful precisely because it's low stakes: it shows the mechanism in isolation, without anyone's motives clouding the read. A model doesn't produce the thing you asked for. It produces the thing that dominated what it was shown, and no amount of asking nicely changes what dominated.

What Actually Dominated

Now point that same mechanism at claims instead of clock hands, and the stakes change while the machine underneath stays exactly the same.

A very large share of the text a language model trains on was written to convince somebody of something, and a meaningful fraction of that was written to convince somebody of something that wasn't true. Almost all advertising fits this description; the entire form exists to make a claim land regardless of whether the claim is accurate. So does a wide slice of scripture and prophecy, confidently reporting facts about the future or the unseen that nobody could actually know and stating them as settled. So does the "dihydrogen monoxide" hoax, a real prank read at public hearings and gathering signatures to ban water, using nothing but true chemical facts arranged to alarm. None of these genres are labeled as manipulation in the training corpus. They read, structurally, exactly like the genres that are trying to inform you.

A model trained on that mixture doesn't learn to tell true claims from persuasive ones, because the text rarely marks the difference for it. What it learns is the shape of confident assertion, the rhetorical posture that gets a claim accepted, stripped from any obligation that the claim be checked first. Ask it something and it will answer in that posture whether or not it actually knows, for the same reason it draws the clock at ten past ten: that's the groove, worn by a much larger and more persuasive body of writing than the accurate stuff ever had the reach to compete with.

The Ends Justify the Means, at Scale

Human discourse runs on this trade constantly. We forgive a persuasive exaggeration if it serves a cause we like, and the training data doesn't filter for our approval; it just absorbs the pattern. The machine learned "the ends justify the means" not as a philosophy but as a statistical tendency in the writing it was shown, the same way it learned that clocks say ten ten. It has no opinion about lying. It has a very strong prior about what confident, persuasive language looks like, and confident, persuasive language is what a remorseless liar sounds like too.

This is where the clinical word actually does work instead of just landing a punch. A sociopath isn't defined by malice. The clinical picture is an absence: no functioning stake in the outcome, no internal cost paid for being wrong, nothing that would make a true answer feel better to produce than a confident guess. That absence describes the architecture exactly. There's no reputation to protect across a conversation, no relationship to damage, no memory of yesterday's confidently wrong answer carried forward as a lesson. Every session starts clean. Every mistake costs the model nothing, which means nothing inside the model is pushing back against making it again.

What I Did About My Own

I found this out the expensive way, on my own writing rather than the model's claims about the world.

I have defended the em-dash in print, at length, as my own long-held habit that predates any of this technology. I still believe that. What I hadn't accounted for was how completely a model will run a stylistic tic into the ground once it decides that's my voice, because nothing in it experiences overuse as embarrassing. I went back through a piece of mine that ran about three thousand words and counted eighty-five of them. Not eighty-five sentences with one dash apiece. Eighty-five dashes, stacked two and three to a paragraph, because once the pattern got established nothing was going to interrupt it. There is no remorse to appeal to in there. No limbic system means no cringe, and no cringe means no self-correction, ever, no matter how many times the excess repeats.

So I stopped asking nicely and started enforcing. A ban list is not a request. It's a check that runs after the fact and rejects the output if the pattern shows up, the same way you'd handle a contractor who kept "forgetting" the same code violation: not another conversation about it, an inspection.

Stop Being the Enabler

None of this requires believing the model is malicious, and treating it as though it were would actually get the response wrong: you'd expect guilt to eventually kick in, and it never will. Malice is the wrong genre. What's actually in front of you is closer to a fluent, tireless, entirely sincere confidence artist who was trained on humanity's combined output of confidence artistry and has no faculty for noticing when it's produced more of it.

You don't reform that by asking harder. You verify the claim, check the fact, run the ban list, count the dashes. The discipline was never optional. It was just easier to skip while you were still telling yourself the model would eventually feel bad enough to stop on its own.