Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do LLMs hallucinate confidently instead of saying 'I don't know'?

The model isn't lying. It was never trained to know when to stop talking.

AI & ML intro Apr 29, 2026 · updated Aug 24, 2026 · 9 min read

On this page

The picture version

The whole answer in six pictures, for a reader who has never thought about how a chatbot picks its next word. The prose below fills in the seams the pictures skip.

1 · The problem

You ask for a source. You get a perfect one that doesn't exist.

you asks for a source what came back a title that sounds right three plausible author names a journal that really exists a year that fits the timeline every part looks correct you paste the title search engine 0 results the paper does not exist
The citation is fluent, specific, and completely fabricated — and nothing about how it was written signals the difference.

2 · Inside the machine

There is no “no such paper” row on the menu.

what you typed, plus what it has written so far the source for that claim is next token? how likely each next token is a plausible title a plausible author a plausible journal a plausible year every option is a token that looks like a citation “no such paper” not on the menu at all
The model has to emit some next token, and the whole menu is made of tokens. Given a prompt shaped like a citation request, the most citation-shaped tokens win.

3 · Two fixes that fail

Teaching the phrase is easy. Teaching the timing is not.

Fix 1: train it to refuse you'd need training text where the right answer is a refusal which questions will this model, after training, get wrong? that label doesn't exist yet the model it describes hasn't been trained yet Fix 2: let humans rate a crisp, specific citation a rater who doesn't check prefers this one “I can't find a source for that” looks like a worse answer so the pressure points toward confident specificity on exactly the questions where confidence isn't warranted
Fix 1 can teach the words, not the moment: “when” is a fact about a model that doesn't exist when the data is collected. Fix 2 can push backwards — to a rater who doesn't check, a fabricated reference reads better than a refusal.

4 · The trap

It isn't unsure. That's the whole problem.

The paper is titled “Attention-Guided one token, overwhelmingly likely sharply peaked — the model is not hedging why it's this sharp the next token is highly constrained by English and by title conventions Confidence about the text is not confidence about the world.
A fabrication doesn't have to look uncertain. Once the model has written “The paper is titled…”, English and title conventions make the next token nearly forced — real paper or not.

5 · The fix that mostly works

Stop asking it to remember. Hand it the document.

your question fetch the real document, pasted into the prompt now read grounded answer reading is something the model is genuinely good at your question fetch retrieval comes back with nothing relevant and then back to picture 2 fluent, ungrounded Grounding narrows the opening. It doesn't close it.
Retrieval and tool calls change the question from recall a source to read this source, and reading is the easier job — but when retrieval returns nothing, the model is free to be fluent from nothing again.

6 · Keep this card

The whole thing on one index card.

hallucination = a next-token machine + no built-in “I don't know” + a prompt it can't answer an improviser who was never told a scene can end
Picture to keep: an improviser on stage who has never been told a scene can end — delivering the next line is the job. Where the picture breaks: an improviser knows they're improvising, and there is no matching “I'm making this up” state inside the model to check.

Why it exists

You’re finishing a document and you need a source for a claim, so you ask a chatbot. Back comes a citation: a paper title that sounds exactly right, three plausible author names, a journal that genuinely exists, a year that fits the timeline. You paste the title into a search engine and get nothing. The paper does not exist.

That fabricated citation is the running example for this post, because the natural next question — why didn’t it just say “I don’t know”? — has an answer that explains almost everything else these systems do. A search engine would have returned zero results. A junior colleague would have said “I’m not sure, let me check.” The model wrote a confident paragraph.

Nothing is broken. An LLM is a machine that takes the text so far and produces, for every possible next token, a number saying how likely that token is to come next. It was trained on one objective: make that spread of numbers match the token that actually came next in a giant pile of human-written text. Every gradient step rewarded producing the most likely continuation. No step explicitly rewarded noticing that the continuation it was about to produce wasn’t grounded in anything.

So when you ask for a source and no such source is encoded in its parameters, the model still has to emit a next token. The distribution has no “no-paper” option; it has tokens. And given a prompt that looks like a citation request, the most probable tokens are the ones that look like a citation. Out comes a fluent one. The fluency is the whole problem — it’s the same fluency that makes the thing useful everywhere else.

Why it matters now

Every product built on an LLM inherits this. A support bot that invents a refund policy, a coding assistant that imports a function that was never in the library, a research tool that fabricates your citation — different surfaces, one cause: the model was asked something it didn’t know, and its job description is “produce fluent text,” not “produce true text.”

It’s also why the engineering response almost never lives inside the model. It lives around it: RAG to put real documents in the prompt, tool calls so the model looks things up instead of recalling them, structured outputs so a validator can reject malformed answers, and human review on anything load-bearing. If you’ve wondered why so much agent infrastructure amounts to “shove ground truth into the prompt at the last possible moment” — this is why.

The short answer

hallucination = next-token model + no built-in "I don't know" + a prompt the model can't actually answer

Picture to keep: an improviser on stage who has never been told a scene can end. Ask for a paper title and you get a paper title, delivered with the same poise as everything else, because delivering the next line is the job. Where that picture breaks: an improviser knows they’re improvising. There’s no corresponding “I’m making this up” state inside the model to check.

The model is doing exactly what it was trained to do — emit a fluent continuation. When the prompt asks for a fact that isn’t reliably encoded in the weights, “fluent” and “true” come apart, and “fluent” wins, because “fluent” is what got optimized.

How it works

The interesting way to see this is to try to fix it and watch each fix fail.

Fix attempt 1: train it to say “I don’t know.” Add examples where the correct answer is an admission of ignorance.

Why it breaks: pretraining’s loss compares the model’s predicted distribution against the token that actually followed in the text. To train “say I don’t know here,” you’d need training text where the right continuation is a refusal — which means knowing, in advance, which questions this particular model, after training will fail on. That label doesn’t exist when the data is collected. Pretraining can teach the phrase; it can’t teach when, because “when” is a fact about a model that doesn’t exist yet. Later training stages can attack it — that’s fix attempt 2 — but they work against the base objective, not with it.

Fix attempt 2: have humans rate answers and prefer the honest ones. This is roughly RLHF: show raters pairs of responses, learn what they prefer, optimize for it.

Why it breaks — and it can break backwards: consider what raters reward. Confident, specific, helpful answers score well. A response that hedges on a question the rater believes has an answer looks like a worse response. So the gradient can point toward confident specificity on exactly the questions where confidence isn’t warranted. Our fake citation is the perfect storm here: to a rater who doesn’t check, a crisp fabricated reference looks better than “I can’t find a source for that.” That preference optimization can push this way is a standard account in the literature; how hard it bites in any current production model isn’t knowable from outside, since post-training details are mostly unpublished.

Fix attempt 3: read the model’s own confidence. Surely a fabrication shows up as a flat, low-probability distribution — so threshold on it and refuse.

Why it breaks: this is the subtle one. Sometimes it does work. Often it doesn’t, because the model can be extremely confident — a sharply peaked distribution — about a token that is false, when the training data made that continuation overwhelmingly likely in that context. Having produced “The paper is titled Attention-Guided…”, the next token is highly constrained by English and by title conventions, regardless of whether the paper exists. Confidence in the output distribution is confidence about the text, not about the world. They coincide only when the training data made them coincide. Whether a model’s internal states carry something more like a truth signal, readable separately from the output probabilities, is an open research question — there are encouraging results and no settled answer for current production models.

Fix attempt 4: stop asking it to remember. Retrieve the real documents and put them in the prompt (RAG), or let it call a search tool and read the result.

Why this one mostly works — and where it stops: the question changes from “recall a source” to “read this source,” and reading is something the model is genuinely good at. But it can still misread a retrieved passage, and if retrieval returns nothing relevant, you’re back to attempt 1 — the model is again free to produce a fluent continuation from nothing. Grounding narrows the opening; it doesn’t close it.

Which is why “just say I don’t know” is harder than it sounds. The model would have to (a) detect that this prompt sits in its unknown-unknown zone, (b) override a strong learned prior that a fluent specific answer is what gets rewarded, and (c) emit a refusal that is itself a plausible continuation. All three are learnable, and refusal training and calibration work aim at exactly them — but each is installed against the grain of the base objective rather than falling out of it, which is why the failure keeps leaking back.

You started with hallucination = next-token model + no "I don't know" + an unanswerable prompt. What did the fake citation add that the formula hides? — the model is a mimic of “what someone who knew the answer would say.” When the training data is thick with such people, the mimicry usually lands on something true. When it isn’t, the mimicry continues anyway, because mimicry is what the weights compute. Nothing about the output changes to signal the difference. That’s not a bug sitting next to the capability; it is the capability, pointed at a question it can’t answer.

Going deeper

A note on what I’m sure of: the mechanism above — next-token objective, no native “don’t know” signal, post-training pressure that can amplify confidence — is well established. The quantitative picture, meaning how often current models fabricate, on which task families, and how much each mitigation buys, moves every few months and varies wildly by benchmark. Treat any specific number you read with the same skepticism you’d apply to a model’s own confident citation.