Why do LLMs hallucinate confidently instead of saying 'I don't know'?
The model isn't lying. It was never trained to know when to stop talking.
On this page
The picture version
The whole answer in six pictures, for a reader who has never thought about how a chatbot picks its next word. The prose below fills in the seams the pictures skip.
1 · The problem
You ask for a source. You get a perfect one that doesn't exist.
2 · Inside the machine
There is no “no such paper” row on the menu.
3 · Two fixes that fail
Teaching the phrase is easy. Teaching the timing is not.
4 · The trap
It isn't unsure. That's the whole problem.
5 · The fix that mostly works
Stop asking it to remember. Hand it the document.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’re finishing a document and you need a source for a claim, so you ask a chatbot. Back comes a citation: a paper title that sounds exactly right, three plausible author names, a journal that genuinely exists, a year that fits the timeline. You paste the title into a search engine and get nothing. The paper does not exist.
That fabricated citation is the running example for this post, because the natural next question — why didn’t it just say “I don’t know”? — has an answer that explains almost everything else these systems do. A search engine would have returned zero results. A junior colleague would have said “I’m not sure, let me check.” The model wrote a confident paragraph.
Nothing is broken. An LLM is a machine that takes the text so far and produces, for every possible next token, a number saying how likely that token is to come next. It was trained on one objective: make that spread of numbers match the token that actually came next in a giant pile of human-written text. Every gradient step rewarded producing the most likely continuation. No step explicitly rewarded noticing that the continuation it was about to produce wasn’t grounded in anything.
So when you ask for a source and no such source is encoded in its parameters, the model still has to emit a next token. The distribution has no “no-paper” option; it has tokens. And given a prompt that looks like a citation request, the most probable tokens are the ones that look like a citation. Out comes a fluent one. The fluency is the whole problem — it’s the same fluency that makes the thing useful everywhere else.
Why it matters now
Every product built on an LLM inherits this. A support bot that invents a refund policy, a coding assistant that imports a function that was never in the library, a research tool that fabricates your citation — different surfaces, one cause: the model was asked something it didn’t know, and its job description is “produce fluent text,” not “produce true text.”
It’s also why the engineering response almost never lives inside the model. It lives around it: RAG to put real documents in the prompt, tool calls so the model looks things up instead of recalling them, structured outputs so a validator can reject malformed answers, and human review on anything load-bearing. If you’ve wondered why so much agent infrastructure amounts to “shove ground truth into the prompt at the last possible moment” — this is why.
The short answer
hallucination = next-token model + no built-in "I don't know" + a prompt the model can't actually answer
Picture to keep: an improviser on stage who has never been told a scene can end. Ask for a paper title and you get a paper title, delivered with the same poise as everything else, because delivering the next line is the job. Where that picture breaks: an improviser knows they’re improvising. There’s no corresponding “I’m making this up” state inside the model to check.
The model is doing exactly what it was trained to do — emit a fluent continuation. When the prompt asks for a fact that isn’t reliably encoded in the weights, “fluent” and “true” come apart, and “fluent” wins, because “fluent” is what got optimized.
How it works
The interesting way to see this is to try to fix it and watch each fix fail.
Fix attempt 1: train it to say “I don’t know.” Add examples where the correct answer is an admission of ignorance.
Why it breaks: pretraining’s loss compares the model’s predicted distribution against the token that actually followed in the text. To train “say I don’t know here,” you’d need training text where the right continuation is a refusal — which means knowing, in advance, which questions this particular model, after training will fail on. That label doesn’t exist when the data is collected. Pretraining can teach the phrase; it can’t teach when, because “when” is a fact about a model that doesn’t exist yet. Later training stages can attack it — that’s fix attempt 2 — but they work against the base objective, not with it.
Fix attempt 2: have humans rate answers and prefer the honest ones. This is roughly RLHF: show raters pairs of responses, learn what they prefer, optimize for it.
Why it breaks — and it can break backwards: consider what raters reward. Confident, specific, helpful answers score well. A response that hedges on a question the rater believes has an answer looks like a worse response. So the gradient can point toward confident specificity on exactly the questions where confidence isn’t warranted. Our fake citation is the perfect storm here: to a rater who doesn’t check, a crisp fabricated reference looks better than “I can’t find a source for that.” That preference optimization can push this way is a standard account in the literature; how hard it bites in any current production model isn’t knowable from outside, since post-training details are mostly unpublished.
Fix attempt 3: read the model’s own confidence. Surely a fabrication shows up as a flat, low-probability distribution — so threshold on it and refuse.
Why it breaks: this is the subtle one. Sometimes it does work. Often it doesn’t, because the model can be extremely confident — a sharply peaked distribution — about a token that is false, when the training data made that continuation overwhelmingly likely in that context. Having produced “The paper is titled Attention-Guided…”, the next token is highly constrained by English and by title conventions, regardless of whether the paper exists. Confidence in the output distribution is confidence about the text, not about the world. They coincide only when the training data made them coincide. Whether a model’s internal states carry something more like a truth signal, readable separately from the output probabilities, is an open research question — there are encouraging results and no settled answer for current production models.
Fix attempt 4: stop asking it to remember. Retrieve the real documents and put them in the prompt (RAG), or let it call a search tool and read the result.
Why this one mostly works — and where it stops: the question changes from “recall a source” to “read this source,” and reading is something the model is genuinely good at. But it can still misread a retrieved passage, and if retrieval returns nothing relevant, you’re back to attempt 1 — the model is again free to produce a fluent continuation from nothing. Grounding narrows the opening; it doesn’t close it.
Which is why “just say I don’t know” is harder than it sounds. The model would have to (a) detect that this prompt sits in its unknown-unknown zone, (b) override a strong learned prior that a fluent specific answer is what gets rewarded, and (c) emit a refusal that is itself a plausible continuation. All three are learnable, and refusal training and calibration work aim at exactly them — but each is installed against the grain of the base objective rather than falling out of it, which is why the failure keeps leaking back.
You started with hallucination = next-token model + no "I don't know" + an unanswerable prompt. What did the fake citation add that the formula hides? —
the model is a mimic of “what someone who knew the answer would say.” When
the training data is thick with such people, the mimicry usually lands on
something true. When it isn’t, the mimicry continues anyway, because mimicry is what the weights
compute. Nothing about the output changes to signal the difference. That’s not a
bug sitting next to the capability; it is the capability, pointed at a question
it can’t answer.
Famous related terms
- Calibration —
calibration ≈ stated confidence matches actual accuracy— a perfectly calibrated model that says “70%” is right 70% of the time. This is exactly what fix attempt 3 was hoping for. - RAG —
RAG = retrieval + LLM generation— the standard industrial answer: hand the model the text so it doesn’t have to recall it. - Tool use / function calling —
tool use = model emits structured calls + harness executes them— replaces “remember the answer” with “go look it up.” - Grounding —
grounding ≈ tying an answer to a real source— the umbrella term; RAG and tool use are two ways to do it. - Confabulation —
confabulation ≈ hallucination, minus any awareness of fabricating— borrowed from neurology; some researchers prefer it because it captures the “fluent invention, no internal flag” flavour better. - Refusal training —
refusal training = post-training that rewards "I don't know" / "I won't"— the direct attack on fix attempt 1. Helps; doesn’t solve.
Going deeper
- Ji et al., Survey of Hallucination in Natural Language Generation (2022) — answers “is this one phenomenon or several?” with a taxonomy that predates the chatbot era and applies beyond LLMs.
- Kadavath et al., Language Models (Mostly) Know What They Know (2022) — the direct empirical attack on fix attempt 3: can a model predict its own correctness? Read it for both the “yes, somewhat” and the caveats.
- Any current frontier model card’s “limitations” section — the rabbit hole for finding out which failure modes the people who built the thing consider worth warning you about, which is usually more candid than the marketing.
A note on what I’m sure of: the mechanism above — next-token objective, no native “don’t know” signal, post-training pressure that can amplify confidence — is well established. The quantitative picture, meaning how often current models fabricate, on which task families, and how much each mitigation buys, moves every few months and varies wildly by benchmark. Treat any specific number you read with the same skepticism you’d apply to a model’s own confident citation.