How can I tell when an LLM is making the answer up?
True answers and fabricated ones come out of the same pipe, in the same tone. There's no red light. But there are seams — places hallucinations cluster, shapes they tend to take, tells you can learn to read.
On this page
The picture version
Six pictures for a reader who has never thought about this before. The prose below fills in the seams the pictures skip.
1 · The problem
True answers and invented ones come out of the same pipe.
2 · The two obvious tests
Both of them read the prose. The prose is the wrong place to look.
3 · The move that works
Fabrication has to land somewhere you can check.
4 · Where to spend the check
Fabrication isn’t spread evenly. It pools where the text ran out.
5 · The honest limit
The bags you never open stay unknown — and some of them are full.
6 · Keep this card
The whole method on one index card.
Why it exists
You’re writing an essay. You ask ChatGPT for a citation, and it gives you a paper title, two authors, a journal, a year. You paste it into your draft. Hours later you go to actually open the paper — and it doesn’t exist. The authors are real. The journal is real. The year is plausible. The paper is not.
That fabricated citation is the running example for this post, and that experience — the answer that looked right, sounded right, and turned out to be confidently invented — is what the post is about. The hard part is that every answer the model gives looks like an answer. True ones and fabricated ones come out of the same pipe, in the same tone, with the same conviction. There is no little red light that flips on when the model is making it up.
That’s the whole problem. The system that wrote the hallucination — fluently emitting a plausible-looking continuation when it didn’t actually know — is the same system that wrote the answer that turned out to be right. From the surface, they’re indistinguishable.
So the practical question, the one a curious user actually has, isn’t “why do hallucinations exist?” — it’s “given that they do, how do I tell, in the moment, whether the thing in front of me is one?”
The honest answer is: you often can’t, not from the text alone. But there are seams. There are places hallucinations cluster, shapes they tend to take, and tells that come from how the model produces them. None of these are decisive. Together, they’re enough to know when to stop trusting the prose and go check.
Why it matters now
LLMs are no longer only toy chatbots. They write code that ships, draft text that gets sent, and are being pushed into regulated settings — legal, medical, financial — often ahead of the evaluation practices around them. The cost of acting on a confident-but-wrong answer scales with what you’re doing with it. Your fabricated citation in a school essay is embarrassing. The same failure inside a deployed agent is a bug. The same failure in a clinical decision-support tool is harm.
This is also why “trust but verify” is too generous a frame. The model is not actually trying to deceive you, but the output is fluent enough that trust is the default state without any conscious decision. The job isn’t “decide whether to trust this.” The job is “notice the specific shapes that should make you stop trusting it, and check the ones that matter.”
The short answer
spotting a hallucination ≈ find the checkable claim + check it + read the rest for tells
Picture to keep: a customs officer who can’t read minds and doesn’t try. They don’t interrogate everyone at the desk — they pull the one bag whose declared weight doesn’t match, and open it. Except that a real officer eventually searches everyone; here the unopened bags stay unopened, and some of them are full. Everything you don’t check remains genuinely unknown.
The checkable claims — names, numbers, citations, code identifiers, dates, quotes — are where fabrication has to land somewhere falsifiable, so they’re where you catch it. The rest of the prose isn’t necessarily true or false in a checkable way; for that, you read for tells.
How it works
The useful thing is to watch the two detectors everyone reaches for first fail, because failing tells you where the real signal lives.
Naive detector #1: read the tone. Surely a made-up answer sounds less sure than a real one. Why it breaks: the model’s prose tone tracks the register of the text it was trained on, not its actual certainty — a citation-shaped answer arrives in the register citations arrive in. Your fabricated citation came out in exactly that register, because it was generated by the same process that generates real ones. Tone is not a reliable channel for the model’s uncertainty.
Naive detector #2: ask it. “Are you sure?” Why it breaks: models trained on human preference data tend to go along with whatever the user pushes for — sycophancy, a measured effect across several assistants (Sharma et al., 2023). They’ll insist they’re sure about something fabricated, or apologize and reverse a correct answer, depending on how your question lands. You’re not reading their certainty; you’re prompting for a continuation in the certainty-or-doubt register. (There is a separate, more interesting question about whether internal model states encode something like usable confidence; see below. The answer is a qualified yes, and it still isn’t in the prose.)
What’s left is the checkable claim. Fabrication has to land somewhere falsifiable — a name, a number, a function that either exists in the library or doesn’t. That’s the bag you open. Everything below is about knowing which bag.
Where they cluster
Long-tail specifics. The model has seen a lot of text on the most
famous topics and very little on the next million, so this is a rule of
thumb about where to look rather than a measured rate. Common,
well-documented things — how TCP works, what git rebase does, how
photosynthesis works — are lower-risk, though not safe. The further
down the long tail you go (the third author on a 2014 paper, an
obscure library’s API, the population of a small town in 1973), the
higher your prior for fabrication should be.
Anything past the training cutoff, asked with no way to look it up. Models are trained on text collected up to some cutoff date, and nothing after it is in there. Versions, releases, current events, prices, who runs what company today. If the model has no tool/search/retrieval available, the prompt still looks like a question it should be able to answer, and a confident guess is a common failure mode. (When the product does have search or RAG attached, this risk is partly absorbed there — see below.)
Citations and quotations. This is the canonical hallucination shape: a paper title that sounds right, plausible authors, a journal that exists, a year that fits, and the contents wrong. Variants include real authors with the wrong title, real titles with the wrong year, real papers cited for claims they don’t make. Direct quotes attributed to specific people are similar. The model has learned the shape of a citation extremely well; the shape doesn’t constrain the contents.
Code that calls things by name. Function names, library imports,
command-line flags, API endpoints. The model knows what kind of
name should appear in this position and produces something that fits
the pattern — df.read_excel_safe(path) looks exactly like a pandas
method, and isn’t one. Same failure as your citation: the shape is
learned, the contents aren’t constrained by it.
Numbers with too many digits. “The 2019 study showed a 37.4% improvement” should make you more suspicious than “the study showed roughly a third improvement.” Specificity is cheap to fabricate and feels authoritative.
Internal, organizational, or proprietary detail. Undocumented APIs, recent policy changes, who reports to whom at a specific company, internal tooling. The model has seen public traces of these and tends to fill the gaps with something that looks plausible.
Shapes the prose takes
Everything in this subsection is heuristic — pattern I and others have noticed, not a measured detection rate. None of it is individually conclusive, since true answers sometimes take these shapes too. Read them as reasons to spend a check, never as verdicts.
Smooth confidence on a question that should be hard. If you know the question genuinely doesn’t have a clean answer, and the model gives you a clean answer with no hedging, look again. Real expertise on hard questions sounds messier than that.
“Famously” / “well-known” / “as everyone knows” doing load-bearing work. These phrases get attached to things that aren’t actually famous when the model has nothing more concrete to say. Fluency filler, not evidence.
Excessive structural neatness. Three reasons, three examples, three counterpoints, all the same length. Real arguments are usually lopsided. A suspiciously balanced list can mean the model filled out a shape rather than reasoning — it can also just mean a tidy answer, so this one is weak on its own.
A specific date with no surrounding context. “In 1987, Smith showed that…” — and Smith is never mentioned anywhere else, before or after. A real expert citing a real result usually has more they could say about it.
Doubling down with new specifics when challenged. If you push back on a claim and the model produces a more elaborate justification with new specific details that weren’t in the original answer, those new details are especially worth checking. Absent a tool call, the model isn’t going off and looking anything up — it’s generating more text that fits the request, and “request” now includes “defend the previous claim.”
Self-contradiction between sessions. Ask the same question twice in different chats. If the answers disagree on specifics, at minimum the model isn’t pulling from a stable source. One could be right and the other wrong; both could be partially confabulated. The disagreement itself is the signal — which one is true is a separate check.
The confidence signal that does exist (and where it isn’t)
Both failed detectors above share an assumption: that the model’s uncertainty is somewhere in the text. It largely isn’t. But that’s a claim about the prose, not about the model.
Ask the question differently — put it in a constrained format and ask the model to score its own answer — and something usable shows up. Kadavath et al. found that on multiple-choice and true/false questions, large models’ stated confidence roughly matched how often they were actually right — reasonably calibrated. Self-evaluation worked better when the model was first shown a few worked examples than with none, and degraded on tasks unlike anything in its training. Note how format-dependent that is — the result is about constrained question formats, not about chat. So “the model has no idea whether it’s right” is too strong. The accurate version is narrower: whatever signal exists is not reliably expressed in free-form prose, which is the only place a normal user ever looks. (The longer treatment is in hallucination: limited privileged access to its own uncertainty, and post-training can attenuate or distort what little surfaces.)
A working procedure
In practice, the checks that actually catch things look like:
- Identify the load-bearing claim. Which specific assertion would, if false, make the answer wrong? It’s usually one or two sentences, not the whole paragraph. If you can’t point to one — if the answer is all soft summary — that’s its own signal.
- Check it against a primary source. Search for the paper, run the code, read the docs, ask a person who’d know. The point isn’t to verify everything; it’s to verify the part that matters. This is also why RAG and tool use exist as product patterns — they replace pure recall with retrieval against a real source. Note that this is grounding, not automatic verification: retrieval can fail, return stale or irrelevant text, or be misread by the model. It moves the failure mode, doesn’t erase it.
- Treat dressed-up specifics as suspicious until checked. If a claim is decorated with a date, a name, and a percentage but you can’t easily check any of them, the right prior is “more likely fabricated than a similar claim phrased loosely.” The veneer of specificity is not evidence.
- Spend more effort in the weak zones. Long-tail, recent, niche, internal, citation-shaped, code-by-name. These are the categories where you should hold the highest prior for fabrication, so this is where a fixed verification budget buys the most.
- For code, run it. The fastest hallucination detector for code is the interpreter. Function-doesn’t-exist errors and import errors are the model telling on itself.
The honest limit
None of this gets you to certain. You can read a hallucinated paragraph that has no tells, written about a topic the model is in its weak zone on, and take it at face value. The only fully reliable answer to “is this a hallucination?” is to verify against ground truth — which means the question of whether to trust an LLM is, in the end, the question of when verification is cheap enough to be worth doing. Where it isn’t, the right move is often not to ask the model in the first place.
You started with spotting a hallucination ≈ find the checkable claim + check it + read the rest for tells. What did watching the two naive
detectors fail add? — + the tells are priors, not evidence. They tell
you where to spend a check, never whether something is true. That’s
why the fabricated citation you opened with was catchable in ten seconds
and the smooth summary paragraph next to it may still be wrong: one had
somewhere to land, the other didn’t.
Famous related terms
- Hallucination —
hallucination = next-token model + no built-in "I don't know" + a prompt the model can't answer. The mechanism this post is the user-facing dual of. - Calibration —
calibration ≈ stated confidence matches actual accuracy. Measured calibration is much better in constrained formats like multiple-choice than in free-form prose. Whether that gap has narrowed for whatever model you’re using today is worth checking rather than assuming — but the prose tone still isn’t where you read certainty. - Sycophancy —
sycophancy = model preferring user-pleasing answer over correct one. A measured failure mode of preference-trained models; one reason asking “are you sure?” is unreliable as a check (the others being poor free-form calibration and prompt sensitivity). - RAG —
RAG = retrieve relevant text + prompt the model with it. The product-level answer to “verify before generating.” - Tool use / function calling — letting the model look things up instead of recalling them. Trades a recall failure for a retrieval failure, which is usually a much better trade — but the model can still misread what it fetched.
- Confabulation — borrowed from neurology, and sometimes offered as a better name, on the grounds that it captures “fluent fabrication with no awareness of fabricating” more accurately than “hallucination.” The argument has been made repeatedly and “hallucination” is still the term in general use.
- Grounding — umbrella term for tying an answer to a real source. The structural fix; the tells in this post are the survival skill until grounding is in place.
Going deeper
- Kadavath et al., Language Models (Mostly) Know What They Know (2022) — the primary source for “does the model have any internal signal about whether it’s right?”, including the formats where the qualified yes stops holding.
- Ji et al., Survey of Hallucination in Natural Language Generation (2022) — the explainer to read if you want the taxonomy behind the clusters above, and how researchers actually measure fabrication rates rather than eyeballing them.
- Try this yourself — the rabbit hole with the best return: pick a topic you genuinely know well, ask a frontier model a hard question in it, and read the answer with an expert’s eye. The first time you watch a smooth paragraph fall apart on close inspection, the rest of this post stops being abstract.
Honest gap: the tells above are heuristics, not measured detectors. How often each one fires on a real hallucination versus a true answer varies by model, by topic, and over time — and published hallucination research measures rates of fabrication, not the hit-rate of surface-level reading cues like these. Treat them as priors.
Which is itself the interesting next thread. If you want to keep pulling: which of these tells has anyone actually measured, and at what false-positive rate? How much does attaching search or RAG move the fabrication rate, in numbers rather than vibes? And does the weak-zone map shift between a raw model and the same model inside a product with tools, system prompts, and a citation UI bolted on? I’d read a good answer to any of those.