The AI DownsideDocumenting AI's downsides

Hallucinations

AI Hallucinations Are Still Not Solved

Every release cycle promises fewer made-up facts. The confident fabrications keep arriving on schedule.

Abstract editorial illustration for “AI Hallucinations Are Still Not Solved”.

With every major model release comes the same reassuring note: hallucinations are down, reliability is up, the fabrication problem is largely behind us. And every release, within days, someone posts a screenshot of the new model inventing a citation, a quote, a case, a statistic or a person with total, serene confidence. The rate improves. The category does not disappear. It is worth understanding why, because the gap between “less often” and “solved” is where the real damage happens.

It is not a bug, which is the uncomfortable part

A hallucination is not a glitch the way a crash is a glitch. Large language models generate text by predicting plausible continuations, and a plausible continuation is not the same thing as a true one. The model has no separate store of verified facts it checks against; it has patterns, and a fabricated citation in exactly the right format is, to the model, an excellent pattern. It is doing precisely what it was built to do. The falsehood and the truth are produced by the identical process, which is why the model is equally confident about both.

The model is not lying, because lying requires knowing the truth. It is producing the most likely-looking answer, and likely-looking is a different target from true.

The failure mode gets worse exactly where you can check least

Hallucination is not evenly distributed, and its distribution is perverse. Models fabricate most readily in precisely the situations where you are least equipped to catch them: obscure topics, niche technical details, specific figures, recent events, and anything at the edge of what was well represented in training. Ask about something popular and well-documented and the answer is usually solid. Ask about something rare — the exact thing you turned to the tool for because you did not know it — and the fabrication rate climbs, while your ability to notice drops to zero. The model is most confident and least reliable in the same dark corners where you have no independent way to tell.

This inverts the trust you would place in a human expert. A knowledgeable person becomes visibly hesitant at the edge of their competence — they hedge, they qualify, they say “I'd want to check that.” The model does the opposite: it maintains identical fluency and confidence whether it is on firm ground or inventing wholesale, offering no tell at the exact moment a tell would matter most. The uniformity of its confidence is not a cosmetic flaw. It is the specific property that makes the fabrications dangerous, because it strips away the single cue humans have always used to calibrate how much to believe.

Confidence is the dangerous ingredient

If these systems hedged — “I think, but I am not sure” — hallucination would be a manageable nuisance. The problem is that the fabrications arrive in the same fluent, assured, well-structured prose as the correct answers. There is no tell. A made-up legal case cites a plausible court and year. An invented statistic sits at a believable number. The interface offers no way to distinguish the two, because the model itself cannot.

This is why hallucination has produced genuine, documented harm rather than just funny screenshots. Lawyers have been sanctioned for filing briefs containing citations to cases that never existed, produced by a chatbot and not checked. That is not a hypothetical; it has happened in real courtrooms, more than once, because the fabrications were persuasive enough to survive a busy professional's glance.

The fixes help, and none of them close the gap

The industry's mitigations are real and worth using, but each has a ceiling:

  • Retrieval — grounding answers in real documents fetched at query time — genuinely reduces fabrication, but the model can still misread, misquote or over-extrapolate from the very sources it was handed.
  • “Reasoning” models that work through problems step by step catch some errors and confidently reason their way into others.
  • Citations in answer engines are only as good as the check you do on them, and a fabricated-but-formatted citation defeats the reader who trusts the format.

All of these lower the rate. None of them change the underlying fact that the system's job is to produce plausible text, and plausible text is sometimes false. You cannot fully suppress a behaviour that is identical, mechanically, to the behaviour you want.

The honesty we are owed

The complaint here is not that the technology hallucinates — that is inherent, and understood. The complaint is the marketing gap. When a release is sold on “dramatically reduced hallucinations,” a reasonable person hears “I can now trust this.” What is actually true is “it will fabricate slightly less often, still with total confidence, still undetectably.” Those are very different messages, and the second one is the one that would keep the sanctioned lawyer out of trouble.

Fluency is being mistaken for competence

Part of why hallucinations do so much damage is a very human bug, not a machine one: we are wired to read fluent, confident, well-organised language as a sign of knowledge. For all of history, someone who could explain a thing clearly and without hesitation usually understood it, because producing fluent expertise required actually having the expertise. Large language models sever that link. They produce the fluency without the understanding, and our instinct to trust the fluency fires anyway. The model exploits a shortcut in human judgement that was reliable right up until a machine learned to fake the surface.

This is why “just be more careful” is weak advice. The failure is not laziness; it is that the single most useful cue humans have for calibrating trust — confident fluency — has been rendered meaningless in this context, and no replacement cue has taken its place. You cannot feel your way to whether an answer is true, because the feeling of truth and the feeling of fabrication are now identical. The only defence is external verification, which is slow, effortful, and exactly the labour the tool promised to save.

The trap of automation complacency

There is a well-documented human tendency that makes all of this worse over time: the more reliable a system usually is, the less we scrutinise it. It is called automation complacency, and it is why people drive into rivers following a satnav. A model that is right ninety-something percent of the time is, perversely, more dangerous than one that is right half the time, because the high base rate lulls you. You check the first ten answers, they are all fine, you stop checking — and the eleventh, the fabricated one, sails straight through the scrutiny you have quietly abandoned.

This means the better these systems get, the more the residual errors matter, because they arrive inside a wall of correctness that has trained you not to look. A world of mostly-right AI is not a world where hallucinations stop mattering; it is a world where they become harder to catch precisely because they are rarer. The improvement in the average case erodes the vigilance you would need for the bad case, which is the specific reason “it hardly ever gets things wrong now” is cold comfort rather than reassurance.

“I don't know” is the feature nobody ships

The single change that would do most to tame hallucination is also the one the incentives fight hardest: a model that reliably says “I don't know” when it doesn't. Humans trust experts partly because good ones admit the edge of their knowledge, and a system that could do the same — hedging where it is uncertain, declining where it is guessing — would let users calibrate exactly where calibration is needed. The technology to express uncertainty is not the barrier. The barrier is that uncertainty demos badly. A model that frequently says “I'm not sure” looks less impressive on stage and in comparisons than one that answers everything with breezy confidence, even when the confidence is unearned.

So the market quietly rewards the wrong trait. Confident-and-sometimes-wrong beats hesitant-and-honest in a side-by-side demo, in a preference leaderboard, in the gut impression of a new user — and the products are tuned, consciously or not, toward the confidence that sells. This is the deep reason hallucination persists beyond its technical roots: honesty about uncertainty is a competitive disadvantage in the way these systems are currently judged. Until buyers start actively prizing a model that knows its limits — and penalising one that bluffs — vendors will keep shipping the smooth, assured voice that fabricates rather than the humble one that hedges. The fix is partly technical, but it is also a matter of what we, collectively, decide to reward. Right now we reward the bluff.

Use the tools. They are useful. But treat every specific, checkable claim — a citation, a date, a quote, a number, a name — as unverified until you have verified it yourself. That is not cynicism; it is the correct operating procedure for a system whose fluency is not evidence of its accuracy. Until a model can reliably say “I do not know” instead of inventing something that looks like knowing, the fabrication problem is not solved. It is merely quieter, which is arguably worse.

Related grievances

All articles →