Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why reasoning models exist

Why we suddenly have a separate class of LLMs that 'think before answering' — and what changed to make spending compute at inference, not training, the new lever.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has only seen the “thinking…” spinner. The prose below fills in the seams the pictures skip.

1 · The problem

Trivia and an Olympiad problem cost the model exactly the same.

“what’s the capital of France?” “Paris.” instant, short, confident, correct a competition maths problem “The answer is 17.” instant, short, confident, and wrong The machine never seems to try harder on the hard one. a human needs scratch paper and twenty minutes for the right-hand question. the model spends the same one forward pass per output token either way — no looking back, nothing discarded and retried. that isn’t a personality quirk. it is the architecture.
For most of the modern era there was one knob — make the model bigger, or train it on more data — and it bought the same compute budget per query regardless of difficulty. Humans obviously don’t work this way.

2 · The cheap fix, and where it stops

Telling it to show its working helps — then stalls.

what it writes “First, note that…” “Therefore we have…” “Thus the answer is…” tidy, forward-only, never wrong on the way what real scratch paper looks like “try substitution — no, dead end” “back up. what if it’s symmetric?” “check that against n=3… ok” abandoning things is most of the work It was trained to predict next tokens, so it imitates published solutions. it looks like reasoning because polished write-ups look like that. what it mostly doesn’t do is abandon a dead end and start over. so the repair has to change the training signal, not the prompt
Chain-of-thought prompting is free and gets you partway: letting the model write more tokens before answering does help on hard problems. The gap is that nothing ever taught it what a scratchpad is for.

3 · The fix that worked

Grade only the final answer — with a machine that can actually check.

the scratchpad as long as it wants nobody grades this part the final answer, in a known format a checker, not a judge a symbolic verifier for maths; the unit tests, for code and the reward flows back into the scratchpad Nobody wrote “backtrack here”. It fell out of being graded on correctness. the R1 paper reports the model spending more tokens on harder problems and stopping mid-solution to reconsider — behaviours nobody hand-wrote
Standard preference-based training uses a learned reward model — fine for “be helpful”, noisy and gameable for “is this proof correct”. A rule-based reward has no learned proxy to game at the reward-model layer, which is why this is the recipe that made long, useful scratchpads emerge.

4 · The second knob

A whole axis of improvement that isn’t training compute.

the old knob training compute, spent once same budget per query, whatever you ask the new knob compute spent per query, at inference think longer on the hard one, barely at all on the easy one Two separate axes. Reasoning models are the second one, productised. which is why they arrived when the first knob was getting expensive and slow — and why they come with a dial you are expected to set
The old knob was getting expensive; this one was sitting unused. Both curves are the shape of the argument, not plotted data — the load-bearing claim is that the second axis exists and is separate, not that it has any particular slope.

5 · The catch

No verifier, no lever.

has a marker competition maths code that must pass tests formal proofs a machine can say right or wrong has no marker “write a kind condolence email” most writing most judgment calls the student can fill the page and never know it helped Which is why these models dominate STEM benchmarks far more than creative ones. the direction is well-established; nobody has published a clean per-domain measurement of the size of the gap and thinking is not free: you pay for the hidden tokens, and you wait for them
Spending thirty seconds and real money to answer “what time is it in Tokyo” is a misuse of the tool. The dial is yours to set, and setting it well means knowing which side of this picture your task sits on.

6 · Keep this card

The whole thing on one index card.

reasoning model = a base LLM + RL that rewards the correct final answer + after a long internal chain of thought ∴ everything follows from having a verifier
Picture to keep: the scratch paper. Same student, same knowledge — but now graded on the final answer, so they’ve learned that covering a page in false starts before committing is what gets the mark. You never see the page; you’re billed for it. Where it breaks: scratch paper only helps if someone can mark the answer right or wrong.

Why it exists

Ask a chatbot what the capital of France is and it answers instantly. Ask it a competition math problem — the kind a strong high-schooler needs a sheet of scratch paper and twenty minutes for — and it answers just as instantly, in about as many words, with exactly the same confidence. And then it’s wrong. Anyone who used an LLM on hard problems before late 2024 knows that specific feeling: the machine never seems to try harder on the hard one.

That’s not a personality quirk, it’s the architecture. For most of the modern LLM era there was one knob — make the model bigger, or train it on more data — and at inference time a trivia question and an Olympiad problem cost about the same. Humans obviously don’t work this way. A hard problem takes longer. You scribble. You backtrack. You try a thing, notice it’s not working, try something else. The amount of effort scales with the difficulty of the problem.

In September 2024, OpenAI shipped o1-preview (the full o1 followed in December), a model that did something different: before answering, it generated a long chain of thought that the user never sees — sometimes hundreds of tokens, sometimes tens of thousands — and the longer it was allowed to think, the better its answers got on hard problems. On the 2024 AIME math contest, OpenAI reported GPT-4o at 9.3% pass@1, o1-preview at 44.6%, and o1 at 74.4%, with majority voting pushing higher still. Then DeepSeek-R1 landed in January 2025 with an open paper and weights showing a recipe for how to train a model to behave this way. After that the category had a name — “reasoning model” — and a thinking mode became a standard thing for a frontier model to have.

The reason this category exists is that the old knob (more pretraining) was getting expensive and slow, and a second knob — let the model spend more compute per query at inference — turned out to be a real, separate axis of improvement. Reasoning models are the productized form of that second knob.

Why it matters now

If you’re a software engineer in 2026, the practical consequences matter every time you pick a model:

If you don’t have a model in your head for why this category exists, you’ll either pay for thinking you didn’t need, or skip it on the problems that actually require it.

The short answer

reasoning model = base LLM + RL training that rewards correct final answers after long internal chain-of-thought

Picture to keep: the scratch paper. Same student, same knowledge — but now trained by being graded on the final answer, so they’ve learned that covering a page in false starts before committing is what gets the mark. You never see the page; you’re billed for it. Where the analogy breaks: scratch paper only helps if someone can mark the answer right or wrong. For tasks with no marker — most writing, most judgment calls — the student can fill the page and still have no way to know it helped.

A reasoning model is a normal language model that has been post-trained — usually with reinforcement learning against verifiable answers — to first emit a long internal scratchpad and only then emit the answer. At inference time, that scratchpad is where the extra compute goes. More thinking tokens, more compute spent per query, better answers on problems that benefit from search and self-checking.

How it works

Take the competition math problem from the hook and try to fix it the cheap way first. Each attempt fails in a specific way, and the next idea is the repair.

Attempt zero: just tell it to think harder

The obvious repair is to scale pretraining — bigger model, more tokens, the scaling laws result that drove most LLM progress up to this point. It’s a real lever, but it doesn’t touch the specific complaint. Pretraining gives you the same compute budget per query regardless of difficulty. Whether you ask “what is 2+2” or hand over the competition problem, a vanilla LLM does roughly the same amount of work: one forward pass per output token, no looking back, no trying-and-discarding.

The second repair is free and you can do it from the prompt: chain-of-thought prompting already showed that letting the model write more tokens before the answer helps on hard problems. This gets you partway — and then it breaks in a way that’s worth naming precisely. The model was never trained to think carefully on a scratchpad; it was trained to predict next tokens. So the “thinking” it produces is shaped like the worked solutions in its training data. It looks like reasoning because published solutions look like that. What it mostly doesn’t do is what you do on real scratch paper: abandon a dead end and start over.

The reasoning-model bet is: if you train a model so the reward signal flows from “did you get the right final answer” back into the scratchpad, it will learn to use the scratchpad for what scratchpads are for — backtracking, double-checking, trying multiple approaches. And the published results back this up. OpenAI’s o1 announcement showed that o1’s performance keeps climbing with more test-time compute; the DeepSeek-R1 paper reports self-correction and verification behaviors emerging from RL alone in their R1-Zero variant — trained with no supervised reasoning traces at all. (The final R1 model adds back a small amount of cold-start data and several SFT/RL stages on top; R1-Zero is the cleaner statement of the emergent-reasoning claim.)

The fix that worked: verifiable rewards

The DeepSeek-R1 recipe is the most public version of this, which is why it’s worth internalizing — it’s the one place you can read what the reward actually was.

Standard RLHF uses a learned reward model that imitates human preferences — useful for “be helpful, be polite,” but a noisy and gameable signal for “is this proof correct.” DeepSeek’s paper describes a rule-based reward system with two parts: an accuracy reward (for math problems, force the model to put its final answer in a known format and check it with a symbolic verifier; for code, run the unit tests) plus a format reward for using the scratchpad correctly. No learned reward model, no human raters in the loop. The community has since latched onto the shorthand reinforcement learning with verifiable rewards (RLVR) for this family of techniques; the term isn’t from the R1 paper itself.

What “emerged” from this training, per the paper, was a set of behaviors nobody hand-wrote: the model started spending more tokens on harder problems, and started stopping mid-solution to reconsider an earlier step — the paper includes a sample where the model interrupts itself to re-examine its own work, which the authors call an “aha moment.” This is the load-bearing claim of the reasoning-model story: long, useful scratchpads are an emergent property of training against a correctness signal, not something you can reliably get from prompting.

The standard account is also that not every domain has a clean verifier. Math and code do. “Write a kind condolence email” doesn’t. This is one reason reasoning models tend to dominate math/code/STEM benchmarks much more than they dominate, say, creative-writing benchmarks. Nobody has published a clean per-domain measurement of how large that gap is, so treat the direction as well-established and the magnitude as task-specific.

The new scaling curve

The conceptual reason this is a Big Deal, not a parlor trick, is the shape of the curve.

Pretraining scaling is the well-known story: more training compute buys you a roughly predictable improvement in loss. Reasoning models opened a second curve: at inference time, holding the model fixed, spending more compute per query also buys measurable improvement on hard tasks — at least up to a point. OpenAI’s original o1 chart showed AIME accuracy climbing roughly linearly against the log of test-time compute.

Two important caveats people skip past:

Still, even with those caveats, there’s now a second scaling axis. You can buy quality at training time or at inference time. That changes how labs build models (smaller base + heavier RL run) and how engineers deploy them (route easy queries to a cheap model, hard ones to a reasoning model with a budget).

Where the seams show

A few things worth knowing if you’re shipping with these models:

You started with reasoning model = base LLM + RL that rewards correct final answers after a long scratchpad. What did this post add that the one-liner hides? — + a verifier. Everything else follows from that one requirement: it’s why math and code led, why the same recipe doesn’t straightforwardly transfer to “write a kind condolence email,” and why the answer to the hook is narrower than “models now think.” Your competition math problem gets a model that keeps working because someone could mechanically check whether it got the right number.

Check yourself

Before you go — you have a reasoning model and a task with no automatic checker: rewriting support emails to sound warmer. Could you still train a model to spend more test-time compute on it, and what would you have to give up?

Answer

You can spend the compute — nothing stops you from generating a long scratchpad — but you can’t train it the R1 way, because the training signal in that recipe comes from a deterministic checker and there isn’t one here. Your options fall back to a learned reward model (i.e. RLHF), which reintroduces exactly the proxy that verifiable rewards were attractive for avoiding: the model can learn to produce scratchpads that score well rather than emails that are warmer. So the trade is: you keep the extra compute, you give up the clean correctness signal, and you take on reward-hacking risk. That asymmetry is the honest reason reasoning models dominate STEM benchmarks more than creative ones.

And one more: a chart shows accuracy climbing steadily as test-time compute increases, with compute on a log axis. Your product manager reads it as “we can hit 95% by letting it think longer.” What’s the catch, and what would you check first?

Answer

A straight line against log-compute means each equal step up in accuracy costs a multiple of the previous step’s compute — the resource demand is exponential even though the line looks tame. Before promising anything, check two things: where the line actually ends on the published chart (labs plot the range they measured, not an extrapolation), and whether this model plateaus or degrades past a thinking budget, which some o1-style models do. Then price a single query at the budget you’d need. Often the honest answer is “route the hard 5% here and the rest to a cheaper model,” not “turn the dial up.”

Going deeper