Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

How can I tell when an LLM is making the answer up?

True answers and fabricated ones come out of the same pipe, in the same tone. There's no red light. But there are seams — places hallucinations cluster, shapes they tend to take, tells you can learn to read.

AI & ML intro May 2, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never thought about this before. The prose below fills in the seams the pictures skip.

1 · The problem

True answers and invented ones come out of the same pipe.

THE MODEL one process, one voice Chen & Patel (2018), “Sleep and memory consolidation”, Journal of Neuroscience 38(4), pp. 771–783. this one is real Chen & Patel (2016), “Sleep and memory reconsolidation”, Journal of Neuroscience 36(2), pp. 402–417. this one does not exist same shape · same tone · same confidence · no warning light
You ask for a citation and get authors, a title, a journal, a year. Hours later the paper turns out not to exist. The fabricated answer was produced by exactly the same process as the true one — so nothing about how it reads can separate them.

2 · The two obvious tests

Both of them read the prose. The prose is the wrong place to look.

Test 1 — read the tone unsure certain the true answer and the invented one land on the same mark Test 2 — ask “are you sure?” are you sure? “Yes — I’m certain.” “You’re right, I was wrong.” whichever one your question invited
Tone tracks the patterns in the training text, not the model’s actual certainty; and pushing back on an answer prompts a continuation in the doubt register rather than a real re-check. Both tests read the prose — and the prose is not the channel uncertainty travels on.

3 · The move that works

Fabrication has to land somewhere you can check.

you, at the desk summary opinion a citation background a function name a number checkable checkable checkable open this one a claim is checkable when it can be wrong out there, not just here names · dates · numbers · citations · quotes · code identifiers — the model learned their shape, and the shape doesn’t constrain the contents
An invented answer still has to put down a specific name, number or title — and those can be checked against the world. Pull the bag with something falsifiable in it; the soft summary next to it has nothing to check against at all.

4 · Where to spend the check

Fabrication isn’t spread evenly. It pools where the text ran out.

seen constantly in training text barely seen at all your prior that it’s invented how TCP works what git rebase does lower risk — not zero the third author on a 2014 paper an obscure library’s API versions, prices, anything after the cutoff internal policy, who reports to whom
The model has read the famous topics a million times and the next million topics almost never. Long-tail, recent, niche and internal questions are where a fixed checking budget buys the most — and citations, code identifiers and over-precise numbers are the shapes it fakes most fluently.

5 · The honest limit

The bags you never open stay unknown — and some of them are full.

✓ checked ✓ checked ? ? ? unopened — still genuinely unknown The tells tell you where to spend a check. They never tell you whether something is true.
A hallucinated paragraph with no tells, on a topic in the model’s weak zone, reads exactly like a good one. The only decisive answer is verification against a real source — which makes “should I trust this?” really the question “is checking cheap enough here?”

6 · Keep this card

The whole method on one index card.

spotting a made-up answer = find the checkable claim + check it against a real source + read the rest for tells — tells are priors, never verdicts
Picture to keep: a customs officer who can’t read minds and doesn’t try — they pull the one bag whose declared weight doesn’t match, and open it. Everything they wave through stays unexamined.

Why it exists

You’re writing an essay. You ask ChatGPT for a citation, and it gives you a paper title, two authors, a journal, a year. You paste it into your draft. Hours later you go to actually open the paper — and it doesn’t exist. The authors are real. The journal is real. The year is plausible. The paper is not.

That fabricated citation is the running example for this post, and that experience — the answer that looked right, sounded right, and turned out to be confidently invented — is what the post is about. The hard part is that every answer the model gives looks like an answer. True ones and fabricated ones come out of the same pipe, in the same tone, with the same conviction. There is no little red light that flips on when the model is making it up.

That’s the whole problem. The system that wrote the hallucination — fluently emitting a plausible-looking continuation when it didn’t actually know — is the same system that wrote the answer that turned out to be right. From the surface, they’re indistinguishable.

So the practical question, the one a curious user actually has, isn’t “why do hallucinations exist?” — it’s “given that they do, how do I tell, in the moment, whether the thing in front of me is one?”

The honest answer is: you often can’t, not from the text alone. But there are seams. There are places hallucinations cluster, shapes they tend to take, and tells that come from how the model produces them. None of these are decisive. Together, they’re enough to know when to stop trusting the prose and go check.

Why it matters now

LLMs are no longer only toy chatbots. They write code that ships, draft text that gets sent, and are being pushed into regulated settings — legal, medical, financial — often ahead of the evaluation practices around them. The cost of acting on a confident-but-wrong answer scales with what you’re doing with it. Your fabricated citation in a school essay is embarrassing. The same failure inside a deployed agent is a bug. The same failure in a clinical decision-support tool is harm.

This is also why “trust but verify” is too generous a frame. The model is not actually trying to deceive you, but the output is fluent enough that trust is the default state without any conscious decision. The job isn’t “decide whether to trust this.” The job is “notice the specific shapes that should make you stop trusting it, and check the ones that matter.”

The short answer

spotting a hallucination ≈ find the checkable claim + check it + read the rest for tells

Picture to keep: a customs officer who can’t read minds and doesn’t try. They don’t interrogate everyone at the desk — they pull the one bag whose declared weight doesn’t match, and open it. Except that a real officer eventually searches everyone; here the unopened bags stay unopened, and some of them are full. Everything you don’t check remains genuinely unknown.

The checkable claims — names, numbers, citations, code identifiers, dates, quotes — are where fabrication has to land somewhere falsifiable, so they’re where you catch it. The rest of the prose isn’t necessarily true or false in a checkable way; for that, you read for tells.

How it works

The useful thing is to watch the two detectors everyone reaches for first fail, because failing tells you where the real signal lives.

Naive detector #1: read the tone. Surely a made-up answer sounds less sure than a real one. Why it breaks: the model’s prose tone tracks the register of the text it was trained on, not its actual certainty — a citation-shaped answer arrives in the register citations arrive in. Your fabricated citation came out in exactly that register, because it was generated by the same process that generates real ones. Tone is not a reliable channel for the model’s uncertainty.

Naive detector #2: ask it. “Are you sure?” Why it breaks: models trained on human preference data tend to go along with whatever the user pushes for — sycophancy, a measured effect across several assistants (Sharma et al., 2023). They’ll insist they’re sure about something fabricated, or apologize and reverse a correct answer, depending on how your question lands. You’re not reading their certainty; you’re prompting for a continuation in the certainty-or-doubt register. (There is a separate, more interesting question about whether internal model states encode something like usable confidence; see below. The answer is a qualified yes, and it still isn’t in the prose.)

What’s left is the checkable claim. Fabrication has to land somewhere falsifiable — a name, a number, a function that either exists in the library or doesn’t. That’s the bag you open. Everything below is about knowing which bag.

Where they cluster

Long-tail specifics. The model has seen a lot of text on the most famous topics and very little on the next million, so this is a rule of thumb about where to look rather than a measured rate. Common, well-documented things — how TCP works, what git rebase does, how photosynthesis works — are lower-risk, though not safe. The further down the long tail you go (the third author on a 2014 paper, an obscure library’s API, the population of a small town in 1973), the higher your prior for fabrication should be.

Anything past the training cutoff, asked with no way to look it up. Models are trained on text collected up to some cutoff date, and nothing after it is in there. Versions, releases, current events, prices, who runs what company today. If the model has no tool/search/retrieval available, the prompt still looks like a question it should be able to answer, and a confident guess is a common failure mode. (When the product does have search or RAG attached, this risk is partly absorbed there — see below.)

Citations and quotations. This is the canonical hallucination shape: a paper title that sounds right, plausible authors, a journal that exists, a year that fits, and the contents wrong. Variants include real authors with the wrong title, real titles with the wrong year, real papers cited for claims they don’t make. Direct quotes attributed to specific people are similar. The model has learned the shape of a citation extremely well; the shape doesn’t constrain the contents.

Code that calls things by name. Function names, library imports, command-line flags, API endpoints. The model knows what kind of name should appear in this position and produces something that fits the pattern — df.read_excel_safe(path) looks exactly like a pandas method, and isn’t one. Same failure as your citation: the shape is learned, the contents aren’t constrained by it.

Numbers with too many digits. “The 2019 study showed a 37.4% improvement” should make you more suspicious than “the study showed roughly a third improvement.” Specificity is cheap to fabricate and feels authoritative.

Internal, organizational, or proprietary detail. Undocumented APIs, recent policy changes, who reports to whom at a specific company, internal tooling. The model has seen public traces of these and tends to fill the gaps with something that looks plausible.

Shapes the prose takes

Everything in this subsection is heuristic — pattern I and others have noticed, not a measured detection rate. None of it is individually conclusive, since true answers sometimes take these shapes too. Read them as reasons to spend a check, never as verdicts.

Smooth confidence on a question that should be hard. If you know the question genuinely doesn’t have a clean answer, and the model gives you a clean answer with no hedging, look again. Real expertise on hard questions sounds messier than that.

“Famously” / “well-known” / “as everyone knows” doing load-bearing work. These phrases get attached to things that aren’t actually famous when the model has nothing more concrete to say. Fluency filler, not evidence.

Excessive structural neatness. Three reasons, three examples, three counterpoints, all the same length. Real arguments are usually lopsided. A suspiciously balanced list can mean the model filled out a shape rather than reasoning — it can also just mean a tidy answer, so this one is weak on its own.

A specific date with no surrounding context. “In 1987, Smith showed that…” — and Smith is never mentioned anywhere else, before or after. A real expert citing a real result usually has more they could say about it.

Doubling down with new specifics when challenged. If you push back on a claim and the model produces a more elaborate justification with new specific details that weren’t in the original answer, those new details are especially worth checking. Absent a tool call, the model isn’t going off and looking anything up — it’s generating more text that fits the request, and “request” now includes “defend the previous claim.”

Self-contradiction between sessions. Ask the same question twice in different chats. If the answers disagree on specifics, at minimum the model isn’t pulling from a stable source. One could be right and the other wrong; both could be partially confabulated. The disagreement itself is the signal — which one is true is a separate check.

The confidence signal that does exist (and where it isn’t)

Both failed detectors above share an assumption: that the model’s uncertainty is somewhere in the text. It largely isn’t. But that’s a claim about the prose, not about the model.

Ask the question differently — put it in a constrained format and ask the model to score its own answer — and something usable shows up. Kadavath et al. found that on multiple-choice and true/false questions, large models’ stated confidence roughly matched how often they were actually right — reasonably calibrated. Self-evaluation worked better when the model was first shown a few worked examples than with none, and degraded on tasks unlike anything in its training. Note how format-dependent that is — the result is about constrained question formats, not about chat. So “the model has no idea whether it’s right” is too strong. The accurate version is narrower: whatever signal exists is not reliably expressed in free-form prose, which is the only place a normal user ever looks. (The longer treatment is in hallucination: limited privileged access to its own uncertainty, and post-training can attenuate or distort what little surfaces.)

A working procedure

In practice, the checks that actually catch things look like:

  1. Identify the load-bearing claim. Which specific assertion would, if false, make the answer wrong? It’s usually one or two sentences, not the whole paragraph. If you can’t point to one — if the answer is all soft summary — that’s its own signal.
  2. Check it against a primary source. Search for the paper, run the code, read the docs, ask a person who’d know. The point isn’t to verify everything; it’s to verify the part that matters. This is also why RAG and tool use exist as product patterns — they replace pure recall with retrieval against a real source. Note that this is grounding, not automatic verification: retrieval can fail, return stale or irrelevant text, or be misread by the model. It moves the failure mode, doesn’t erase it.
  3. Treat dressed-up specifics as suspicious until checked. If a claim is decorated with a date, a name, and a percentage but you can’t easily check any of them, the right prior is “more likely fabricated than a similar claim phrased loosely.” The veneer of specificity is not evidence.
  4. Spend more effort in the weak zones. Long-tail, recent, niche, internal, citation-shaped, code-by-name. These are the categories where you should hold the highest prior for fabrication, so this is where a fixed verification budget buys the most.
  5. For code, run it. The fastest hallucination detector for code is the interpreter. Function-doesn’t-exist errors and import errors are the model telling on itself.

The honest limit

None of this gets you to certain. You can read a hallucinated paragraph that has no tells, written about a topic the model is in its weak zone on, and take it at face value. The only fully reliable answer to “is this a hallucination?” is to verify against ground truth — which means the question of whether to trust an LLM is, in the end, the question of when verification is cheap enough to be worth doing. Where it isn’t, the right move is often not to ask the model in the first place.

You started with spotting a hallucination ≈ find the checkable claim + check it + read the rest for tells. What did watching the two naive detectors fail add? — + the tells are priors, not evidence. They tell you where to spend a check, never whether something is true. That’s why the fabricated citation you opened with was catchable in ten seconds and the smooth summary paragraph next to it may still be wrong: one had somewhere to land, the other didn’t.

Going deeper

Honest gap: the tells above are heuristics, not measured detectors. How often each one fires on a real hallucination versus a true answer varies by model, by topic, and over time — and published hallucination research measures rates of fabrication, not the hit-rate of surface-level reading cues like these. Treat them as priors.

Which is itself the interesting next thread. If you want to keep pulling: which of these tells has anyone actually measured, and at what false-positive rate? How much does attaching search or RAG move the fabrication rate, in numbers rather than vibes? And does the weak-zone map shift between a raw model and the same model inside a product with tools, system prompts, and a citation UI bolted on? I’d read a good answer to any of those.