Why does chain-of-thought prompting work?
Adding 'let's think step by step' to a prompt makes models measurably better at hard problems. Nobody fully agrees on why, and the wrong story will mislead you about how to use it.
On this page
The picture version
The whole idea in six pictures, for a reader who has never thought about how a chatbot gets an answer. The prose below fills in the seams the pictures skip.
1 · The problem
Same model, same question — two different answers.
2 · Why the fast answer failed
One word out means one trip through the machine.
3 · The key idea
Let it write, and every word buys another trip.
4 · What the mechanism hides
The working isn’t a report of the thinking. It is the paper.
5 · The seam
Nothing checks the paper, and a tidy wrong line reads best of all.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’re repainting a room and you ask a chatbot: four walls, each 4 m by 2.5 m, two coats, and a tin covers 12 m² — how many tins? It answers “5 tins” in a confident half-second. You frown, type “show your working,” and the same model — same weights, same conversation — walks through 40 m², doubled to 80 m², divided by 12, rounds up, and lands on 7. Which is right.
You changed four words of prompt. But notice which four: you didn’t supply a fact, a formula, or a correction. You gave it permission to write more before committing.
That’s an uncomfortable fact if your mental model of a language model is “a next-token predictor”: under that account the content of your prompt should matter, but the length of the model’s own scratchwork shouldn’t. Yet here’s a handle you can pull at prompt time, with nothing but words. Chain-of-thought prompting exists because that handle turned out to be too cheap and too effective to ignore — and because it forced a correction to the mental model. How well one of these systems reasons isn’t a property of the model alone. It’s a property of the model plus how much room you gave it to compute on the way to the answer. We’ll keep the paint tins in view the whole way down.
Why it matters now
The trick stopped being a prompting tip and became a product tier. Reasoning models — the OpenAI o-series, DeepSeek-R1, Anthropic’s extended-thinking modes, Google’s Gemini Thinking variants — all spend an extended run of reasoning tokens before the visible answer. What you get to see of that run differs by vendor: some show the trace, some show only a summary, some hide it entirely. How each one was trained to do it varies and is only partly public; DeepSeek-R1 is the outlier that published its recipe.
Three places this shows up in your week:
- The bill. Providers generally meter the thinking tokens, so a model that thinks for 10k tokens costs you roughly 10k tokens more — check your provider’s pricing page, since the naming and rates differ. If you have a “thinking” toggle, it’s a spend dial.
- Agents. A tool-use loop is chain-of-thought with the scratchpad replaced by real actions: write a thought, call a tool, read the result, write another. The same intermediate-token machinery drives it — see agent harness.
- Benchmark numbers you read. A score on a math set means little without knowing the thinking budget. “Model A beats model B” and “model A was allowed to think five times longer” are easy to confuse.
So the engineer’s question isn’t “should I append think step by step?” It’s “how much intermediate computation does this task actually need, and am I paying for it deliberately?”
The short answer
chain-of-thought = let the model write intermediate tokens before the answer + condition each new token on all the previous ones
Picture to keep: someone doing long division on paper versus in their head. Same person, same arithmetic ability — but the one with paper can park “80” somewhere it doesn’t have to be held. The generated tokens are the paper. Where the analogy breaks: a person can glance back at the paper and notice a mistake, because their reading and their arithmetic are separate faculties. The model’s “reading” of its own paper is done by the same process that wrote it, so a wrong line looks exactly as convincing as a right one.
A transformer does a bounded amount of work per token it produces. Demand the answer as the very next token and you’ve allowed it exactly one pass. Let it write a paragraph of working first and you’ve allowed it hundreds, each one able to read everything written so far. More tokens before the answer means more compute and more externalized intermediate state.
That’s the mechanical story, and it’s solid. The harder question — why the scratch space helps as much as it does — is not fully settled.
How it works
Start with the failing version and let each fix create the next problem.
Attempt 1: just ask for the number. “How many tins?” → “5.” The model gets one forward pass in which to multiply, double, divide, and round up.
Why it breaks: computation per token is bounded by depth — a fixed stack of layers, run once. The paint problem needs several results composed in sequence, and there is nowhere to put the first one while computing the second. There’s a line of theory formalizing exactly this; see Merrill and Sabharwal, The Expressive Power of Transformers with Chain of Thought, which characterizes how the number of intermediate steps changes what a transformer can compute. I’d point you at the paper rather than paraphrase its theorem statements, which are more specific than the intuition here.
Fix: let it write first. Ask for the working, and each emitted token gets its own forward pass. Compute per answer is now something you control.
But more passes only help if later passes can use earlier ones. Where does the intermediate value live?
Fix: nowhere special — it lives in the text. Transformers are autoregressive: token n is conditioned on tokens 1…n−1, including everything the model itself just wrote. When it writes “4 × 2.5 = 10 m² per wall, × 4 walls = 40 m²,” the number 40 is now in the context, as available as if you had typed it. The model wrote itself a note. This is the part worth internalizing: an LLM has no private scratchpad alongside the text. Implementations do carry state between steps (that’s what a KV cache is), but that state is computed from the tokens so far — there’s no side channel where the model can stash a number it didn’t write down.
But notes get written whether or not they’re correct, and nothing external checks them. Suppose the model writes “40 m², two coats, so 60 m².” Every later step conditions on 60. It divides by 12, gets exactly 5, and reports 5 tins with more confidence than before, because the working looks tidy. Models do sometimes catch themselves mid-trace and back up — but nothing in the mechanism forces that, and a wrong premise that reads fluently is the easiest thing in the world to keep building on.
But why should a mere phrase — “show your working” — flip the behavior at all, with no weight change?
The leading candidate explanation: the training corpus. Textbook solutions, math forum threads, commented code, worked examples — the internet is saturated with “here’s the problem, here’s the working, here’s the answer.” The hypothesis is that a model trained on that has absorbed an association between visible working and correct conclusions, so prompting for steps steers it toward the region of its learned distribution where that pattern lives. I find this convincing and it’s widely repeated, but the size of its contribution relative to the extra-compute story is not something the literature has pinned down. Treat it as a hypothesis with good motivation, not a measured share.
Wei et al.’s Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) established the effect with worked examples in the prompt; Kojima et al.’s Large Language Models are Zero-Shot Reasoners (2022) showed the bare trigger phrase “Let’s think step by step” does much of the same job. Those are the canonical empirical references. The literature on why is messier.
But a trick that depends on the user remembering to ask is fragile.
Fix: move it from prompt time to training time. The publicly documented approach — DeepSeek-R1’s — generates many chains per problem, scores the final answers against rewards you can check automatically (math results, code that runs), and reinforces the model toward the chains that led to correct ones. That loop is the core of it; the shipped R1 wraps it in further training stages. Whether o1 and the extended-thinking modes work the same way, I can’t tell you: their recipes aren’t published, and even the R1 report leaves gaps around data mixture and reward shaping. What’s observable from outside is the behavior, not the training.
But now the trace looks like an explanation, and it’s tempting to read it as one. This is the seam that matters most:
- The chain is not a window into the model’s real reasoning. It’s a generated artifact, sampled from the same model. Turpin et al. (2023) showed models producing fluent chains that don’t mention the prompt bias actually driving their answer. Treat a trace as suggestive, not as evidence.
- It can hurt on easy tasks. Forcing a ramble through a trivial classification introduces failure modes a one-token answer didn’t have.
- The benefit isn’t uniform across model sizes. The original papers reported small models benefiting little or getting worse while large ones gained a lot — the “emergence” framing. Whether that’s a real discontinuity or an artifact of the metric has been argued both ways (Schaeffer et al., Are Emergent Abilities of Large Language Models a Mirage?); I don’t think it’s settled.
So, back to the tin of paint. Why was the one-shot answer wrong, when the model
clearly could do the arithmetic? You started with
chain-of-thought = intermediate tokens + each conditioned on the last. What
does the paint example add that the definition hides? — the written tokens
aren’t a transcript of the thinking, they’re where the thinking is kept. One
token of output is one pass with nowhere to put 40 m². That’s why “show your
working” helped, and it’s the same reason a written-down “60 m²” is so hard to
come back from.
Check yourself
Before you go — you switch a customer-support classifier (“is this ticket about billing, shipping, or returns?”) from direct answers to chain-of-thought, and accuracy drops slightly while cost triples. What happened, and does this contradict the mechanism above?
Answer
No contradiction. The mechanism says extra tokens buy extra compute and extra state — but this task never needed composed sub-results; one forward pass was already enough. So you paid for compute that had nothing to do, and you added a new failure surface: the model can now talk itself into a wrong category by generating a plausible-sounding justification and then conditioning on it. Extra scratchpad tokens pay off mainly when the bottleneck was serial computation; when the bottleneck is recognition, you’re mostly buying risk and cost. (Not a law — step-by-step framing can still help a classifier by forcing it to consider the evidence. The point is that it’s no longer free, so measure rather than assume.)
And one more — a colleague argues their agent is trustworthy because its reasoning trace is always sensible and you can read it. What’s wrong with the argument?
Answer
It assumes the trace caused the answer. The trace is sampled from the same model as the answer, so a model swayed by something it never mentions — a leading phrase in the prompt, an ordering bias — will happily produce a clean trace that omits the actual driver. That’s what Turpin et al. showed for chain-of-thought prompting; an agent stack inherits the problem. A readable trace is evidence about what a plausible justification looks like, not evidence about the computation. If you want trustworthiness, verify the output against something external (run the code, check the arithmetic, cite-check the claim) rather than grading the prose.
Famous related terms
- Self-consistency —
self-consistency = sample N chains + majority vote on the final answer— a cheap variance reducer that directly targets the “one bad step poisons everything” failure. - Tree-of-thoughts —
ToT ≈ CoT + branching + a search procedure— explore several paths and prune; more expensive, sometimes much better on planning-shaped tasks. - Reasoning model —
reasoning model = base LLM + training that makes it work at length before answering— the production form of the trick. - RLHF — the older “make the model behave” lever; reasoning training is its sibling — same machinery, different reward signal.
- In-context learning — the broader phenomenon CoT sits inside: behavior adapts to patterns in the prompt with no weight update. See in-context learning.
- Test-time compute —
test-time compute = work spent per query at inference, not at training— CoT is the original way to spend it; sampling, search, and verifier-guided decoding are newer ways.
Going deeper
- Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022) — the primary source for how big the effect is and on which benchmarks, which is the thing most secondhand summaries get vague about.
- Turpin et al., Language Models Don’t Always Say What They Think (2023) — read this to find out how much you can trust a reasoning trace as an explanation. (Short answer: less than you’d like.)
- Lilian Weng, Prompt Engineering — the explainer to read if you want CoT placed in context alongside few-shot prompting, self-consistency, and the rest of the family, with the citations attached.
- DeepSeek-AI, DeepSeek-R1 technical report (2025) — the rabbit hole, for the reader who wants to know what it actually takes to train the prompting trick into the weights.
What I’m confident about: the mechanical story (more tokens = more forward passes = more externalized state) and the empirical effect on hard multi-step tasks. What I’m not: the precise mix of “extra compute,” “self-conditioning,” and “matching a training-distribution pattern” that explains how much CoT helps on a given task. If someone has that decomposition pinned down, ask for the citation.