Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does chain-of-thought prompting work?

Adding 'let's think step by step' to a prompt makes models measurably better at hard problems. Nobody fully agrees on why, and the wrong story will mislead you about how to use it.

AI & ML intermediate Apr 29, 2026 · updated Aug 24, 2026 · 12 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never thought about how a chatbot gets an answer. The prose below fills in the seams the pictures skip.

1 · The problem

Same model, same question — two different answers.

four walls, 4 m × 2.5 m each · two coats · a tin covers 12 m² you ask for the number “how many tins?” answers at once 5 wrong you ask it to show its working “show your working” writes first, answers after 7 right
Four words of prompt changed the answer. No new fact, no correction, no change to the model — so whatever fixed this, it wasn’t knowledge the model was missing.

2 · Why the fast answer failed

One word out means one trip through the machine.

a fixed stack of layers — run once layer layer layer layer out comes one word 5 but the paint problem is four steps, in order 4 × 2.5 × 4 walls = 40 m² × 2 coats = 80 m² ÷ 12 = 6.67 round up = 7 each step needs the one above it — and there is nowhere to park 40 m² four steps, one trip. Something gets dropped.
A model does a bounded amount of work per word it produces. Demand the answer as the very next word and you have allowed it exactly one pass to do all four steps at once.

3 · The key idea

Let it write, and every word buys another trip.

one word out 1 trip 5 all the work squeezed here a paragraph of working out — one trip per word … more words before the answer = more trips through the machine how much it computes is now something you set.
Every word the model emits gets its own full run through the layers. The amount of computing spent on your answer stops being fixed and becomes a dial you turn with the prompt.

4 · What the mechanism hides

The working isn’t a report of the thinking. It is the paper.

WHAT IT HAS WRITTEN SO FAR 4 × 2.5 = 10 m² per wall × 4 walls = 40 m² two coats → 80 m² 80 ÷ 12 = 6.67 read back writing the next word, it can see all of that a private notepad the model writes on instead does not exist — if it isn’t written, it isn’t kept The 40 m² survives because it was typed out, not because it was remembered.
There is no side channel where the model stashes a number it didn’t write down. The visible working is the storage — that is why writing more helps at all.

5 · The seam

Nothing checks the paper, and a tidy wrong line reads best of all.

A CLEAN-LOOKING TRACE walls = 40 m² two coats, so 60 m² ← wrong 60 ÷ 12 = 5 ∴ 5 tins every later step builds on it so it reports 5 same wrong answer — now with working attached, so it looks more convincing The model reads its own paper with the same faculty that wrote it.
Models do sometimes catch themselves and back up, but nothing in the mechanism forces it. A trace is a generated artifact, not a window into the reasoning — read it as suggestive, never as evidence.

6 · Keep this card

The whole thing on one index card.

chain of thought = let it write working before the answer + every new word reads everything written so far the working is where the thinking is kept
Picture to keep: someone doing long division on paper versus in their head — same person, same arithmetic, but the one with paper can park “80” somewhere it doesn’t have to be held. The generated words are the paper.

Why it exists

You’re repainting a room and you ask a chatbot: four walls, each 4 m by 2.5 m, two coats, and a tin covers 12 m² — how many tins? It answers “5 tins” in a confident half-second. You frown, type “show your working,” and the same model — same weights, same conversation — walks through 40 m², doubled to 80 m², divided by 12, rounds up, and lands on 7. Which is right.

You changed four words of prompt. But notice which four: you didn’t supply a fact, a formula, or a correction. You gave it permission to write more before committing.

That’s an uncomfortable fact if your mental model of a language model is “a next-token predictor”: under that account the content of your prompt should matter, but the length of the model’s own scratchwork shouldn’t. Yet here’s a handle you can pull at prompt time, with nothing but words. Chain-of-thought prompting exists because that handle turned out to be too cheap and too effective to ignore — and because it forced a correction to the mental model. How well one of these systems reasons isn’t a property of the model alone. It’s a property of the model plus how much room you gave it to compute on the way to the answer. We’ll keep the paint tins in view the whole way down.

Why it matters now

The trick stopped being a prompting tip and became a product tier. Reasoning models — the OpenAI o-series, DeepSeek-R1, Anthropic’s extended-thinking modes, Google’s Gemini Thinking variants — all spend an extended run of reasoning tokens before the visible answer. What you get to see of that run differs by vendor: some show the trace, some show only a summary, some hide it entirely. How each one was trained to do it varies and is only partly public; DeepSeek-R1 is the outlier that published its recipe.

Three places this shows up in your week:

So the engineer’s question isn’t “should I append think step by step?” It’s “how much intermediate computation does this task actually need, and am I paying for it deliberately?”

The short answer

chain-of-thought = let the model write intermediate tokens before the answer + condition each new token on all the previous ones

Picture to keep: someone doing long division on paper versus in their head. Same person, same arithmetic ability — but the one with paper can park “80” somewhere it doesn’t have to be held. The generated tokens are the paper. Where the analogy breaks: a person can glance back at the paper and notice a mistake, because their reading and their arithmetic are separate faculties. The model’s “reading” of its own paper is done by the same process that wrote it, so a wrong line looks exactly as convincing as a right one.

A transformer does a bounded amount of work per token it produces. Demand the answer as the very next token and you’ve allowed it exactly one pass. Let it write a paragraph of working first and you’ve allowed it hundreds, each one able to read everything written so far. More tokens before the answer means more compute and more externalized intermediate state.

That’s the mechanical story, and it’s solid. The harder question — why the scratch space helps as much as it does — is not fully settled.

How it works

Start with the failing version and let each fix create the next problem.

Attempt 1: just ask for the number. “How many tins?” → “5.” The model gets one forward pass in which to multiply, double, divide, and round up.

Why it breaks: computation per token is bounded by depth — a fixed stack of layers, run once. The paint problem needs several results composed in sequence, and there is nowhere to put the first one while computing the second. There’s a line of theory formalizing exactly this; see Merrill and Sabharwal, The Expressive Power of Transformers with Chain of Thought, which characterizes how the number of intermediate steps changes what a transformer can compute. I’d point you at the paper rather than paraphrase its theorem statements, which are more specific than the intuition here.

Fix: let it write first. Ask for the working, and each emitted token gets its own forward pass. Compute per answer is now something you control.

But more passes only help if later passes can use earlier ones. Where does the intermediate value live?

Fix: nowhere special — it lives in the text. Transformers are autoregressive: token n is conditioned on tokens 1…n−1, including everything the model itself just wrote. When it writes “4 × 2.5 = 10 m² per wall, × 4 walls = 40 m²,” the number 40 is now in the context, as available as if you had typed it. The model wrote itself a note. This is the part worth internalizing: an LLM has no private scratchpad alongside the text. Implementations do carry state between steps (that’s what a KV cache is), but that state is computed from the tokens so far — there’s no side channel where the model can stash a number it didn’t write down.

But notes get written whether or not they’re correct, and nothing external checks them. Suppose the model writes “40 m², two coats, so 60 m².” Every later step conditions on 60. It divides by 12, gets exactly 5, and reports 5 tins with more confidence than before, because the working looks tidy. Models do sometimes catch themselves mid-trace and back up — but nothing in the mechanism forces that, and a wrong premise that reads fluently is the easiest thing in the world to keep building on.

But why should a mere phrase — “show your working” — flip the behavior at all, with no weight change?

The leading candidate explanation: the training corpus. Textbook solutions, math forum threads, commented code, worked examples — the internet is saturated with “here’s the problem, here’s the working, here’s the answer.” The hypothesis is that a model trained on that has absorbed an association between visible working and correct conclusions, so prompting for steps steers it toward the region of its learned distribution where that pattern lives. I find this convincing and it’s widely repeated, but the size of its contribution relative to the extra-compute story is not something the literature has pinned down. Treat it as a hypothesis with good motivation, not a measured share.

Wei et al.’s Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) established the effect with worked examples in the prompt; Kojima et al.’s Large Language Models are Zero-Shot Reasoners (2022) showed the bare trigger phrase “Let’s think step by step” does much of the same job. Those are the canonical empirical references. The literature on why is messier.

But a trick that depends on the user remembering to ask is fragile.

Fix: move it from prompt time to training time. The publicly documented approach — DeepSeek-R1’s — generates many chains per problem, scores the final answers against rewards you can check automatically (math results, code that runs), and reinforces the model toward the chains that led to correct ones. That loop is the core of it; the shipped R1 wraps it in further training stages. Whether o1 and the extended-thinking modes work the same way, I can’t tell you: their recipes aren’t published, and even the R1 report leaves gaps around data mixture and reward shaping. What’s observable from outside is the behavior, not the training.

But now the trace looks like an explanation, and it’s tempting to read it as one. This is the seam that matters most:

So, back to the tin of paint. Why was the one-shot answer wrong, when the model clearly could do the arithmetic? You started with chain-of-thought = intermediate tokens + each conditioned on the last. What does the paint example add that the definition hides? — the written tokens aren’t a transcript of the thinking, they’re where the thinking is kept. One token of output is one pass with nowhere to put 40 m². That’s why “show your working” helped, and it’s the same reason a written-down “60 m²” is so hard to come back from.

Check yourself

Before you go — you switch a customer-support classifier (“is this ticket about billing, shipping, or returns?”) from direct answers to chain-of-thought, and accuracy drops slightly while cost triples. What happened, and does this contradict the mechanism above?

Answer

No contradiction. The mechanism says extra tokens buy extra compute and extra state — but this task never needed composed sub-results; one forward pass was already enough. So you paid for compute that had nothing to do, and you added a new failure surface: the model can now talk itself into a wrong category by generating a plausible-sounding justification and then conditioning on it. Extra scratchpad tokens pay off mainly when the bottleneck was serial computation; when the bottleneck is recognition, you’re mostly buying risk and cost. (Not a law — step-by-step framing can still help a classifier by forcing it to consider the evidence. The point is that it’s no longer free, so measure rather than assume.)

And one more — a colleague argues their agent is trustworthy because its reasoning trace is always sensible and you can read it. What’s wrong with the argument?

Answer

It assumes the trace caused the answer. The trace is sampled from the same model as the answer, so a model swayed by something it never mentions — a leading phrase in the prompt, an ordering bias — will happily produce a clean trace that omits the actual driver. That’s what Turpin et al. showed for chain-of-thought prompting; an agent stack inherits the problem. A readable trace is evidence about what a plausible justification looks like, not evidence about the computation. If you want trustworthiness, verify the output against something external (run the code, check the arithmetic, cite-check the claim) rather than grading the prose.

Going deeper

What I’m confident about: the mechanical story (more tokens = more forward passes = more externalized state) and the empirical effect on hard multi-step tasks. What I’m not: the precise mix of “extra compute,” “self-conditioning,” and “matching a training-distribution pattern” that explains how much CoT helps on a given task. If someone has that decomposition pinned down, ask for the citation.