Why reasoning models exist
Why we suddenly have a separate class of LLMs that 'think before answering' — and what changed to make spending compute at inference, not training, the new lever.
On this page
The picture version
Six pictures for a reader who has only seen the “thinking…” spinner. The prose below fills in the seams the pictures skip.
1 · The problem
Trivia and an Olympiad problem cost the model exactly the same.
2 · The cheap fix, and where it stops
Telling it to show its working helps — then stalls.
3 · The fix that worked
Grade only the final answer — with a machine that can actually check.
4 · The second knob
A whole axis of improvement that isn’t training compute.
5 · The catch
No verifier, no lever.
6 · Keep this card
The whole thing on one index card.
Why it exists
Ask a chatbot what the capital of France is and it answers instantly. Ask it a competition math problem — the kind a strong high-schooler needs a sheet of scratch paper and twenty minutes for — and it answers just as instantly, in about as many words, with exactly the same confidence. And then it’s wrong. Anyone who used an LLM on hard problems before late 2024 knows that specific feeling: the machine never seems to try harder on the hard one.
That’s not a personality quirk, it’s the architecture. For most of the modern LLM era there was one knob — make the model bigger, or train it on more data — and at inference time a trivia question and an Olympiad problem cost about the same. Humans obviously don’t work this way. A hard problem takes longer. You scribble. You backtrack. You try a thing, notice it’s not working, try something else. The amount of effort scales with the difficulty of the problem.
In September 2024, OpenAI shipped o1-preview (the full o1 followed in December), a model that did something different: before answering, it generated a long chain of thought that the user never sees — sometimes hundreds of tokens, sometimes tens of thousands — and the longer it was allowed to think, the better its answers got on hard problems. On the 2024 AIME math contest, OpenAI reported GPT-4o at 9.3% pass@1, o1-preview at 44.6%, and o1 at 74.4%, with majority voting pushing higher still. Then DeepSeek-R1 landed in January 2025 with an open paper and weights showing a recipe for how to train a model to behave this way. After that the category had a name — “reasoning model” — and a thinking mode became a standard thing for a frontier model to have.
The reason this category exists is that the old knob (more pretraining) was getting expensive and slow, and a second knob — let the model spend more compute per query at inference — turned out to be a real, separate axis of improvement. Reasoning models are the productized form of that second knob.
Why it matters now
If you’re a software engineer in 2026, the practical consequences matter every time you pick a model:
- Two pricing modes, not one. Providers that hide the scratchpad still bill you for it: OpenAI counts reasoning tokens as output tokens in API usage. A “thinking” call can cost much more than the same prompt on a non-reasoning model — how much more depends on the provider, the model, and how long it decides to think. Sometimes that’s worth it; sometimes it isn’t. You need a feel for which.
- Latency is now a deliberate trade. Non-reasoning models stream the answer in seconds. A reasoning model can sit silent for tens of seconds — sometimes longer on hard problems — before the first user-visible token. Products typically surface this with a “thinking…” indicator because the silence otherwise looks broken.
- A new dial in your code. Reasoning APIs expose some control over how much the model thinks — a named effort level, a token budget, or both. The parameter names and the allowed values differ by provider, by API surface, and by model generation, so read the docs for the model you’re calling rather than trusting a snippet. Picking a setting well requires understanding what it actually buys.
- Different failure modes. A reasoning model that gets the answer wrong doesn’t fail like a non-reasoning model. It fails after a long, confident-looking deliberation. The hallucination is now reasoned-into.
- Agents lean on them. Long-horizon tool-using agents are unusually sensitive to per-step quality. A reasoning model is often the cheapest way to buy that quality without a fine-tune.
If you don’t have a model in your head for why this category exists, you’ll either pay for thinking you didn’t need, or skip it on the problems that actually require it.
The short answer
reasoning model = base LLM + RL training that rewards correct final answers after long internal chain-of-thought
Picture to keep: the scratch paper. Same student, same knowledge — but now trained by being graded on the final answer, so they’ve learned that covering a page in false starts before committing is what gets the mark. You never see the page; you’re billed for it. Where the analogy breaks: scratch paper only helps if someone can mark the answer right or wrong. For tasks with no marker — most writing, most judgment calls — the student can fill the page and still have no way to know it helped.
A reasoning model is a normal language model that has been post-trained — usually with reinforcement learning against verifiable answers — to first emit a long internal scratchpad and only then emit the answer. At inference time, that scratchpad is where the extra compute goes. More thinking tokens, more compute spent per query, better answers on problems that benefit from search and self-checking.
How it works
Take the competition math problem from the hook and try to fix it the cheap way first. Each attempt fails in a specific way, and the next idea is the repair.
Attempt zero: just tell it to think harder
The obvious repair is to scale pretraining — bigger model, more tokens, the scaling laws result that drove most LLM progress up to this point. It’s a real lever, but it doesn’t touch the specific complaint. Pretraining gives you the same compute budget per query regardless of difficulty. Whether you ask “what is 2+2” or hand over the competition problem, a vanilla LLM does roughly the same amount of work: one forward pass per output token, no looking back, no trying-and-discarding.
The second repair is free and you can do it from the prompt: chain-of-thought prompting already showed that letting the model write more tokens before the answer helps on hard problems. This gets you partway — and then it breaks in a way that’s worth naming precisely. The model was never trained to think carefully on a scratchpad; it was trained to predict next tokens. So the “thinking” it produces is shaped like the worked solutions in its training data. It looks like reasoning because published solutions look like that. What it mostly doesn’t do is what you do on real scratch paper: abandon a dead end and start over.
The reasoning-model bet is: if you train a model so the reward signal flows from “did you get the right final answer” back into the scratchpad, it will learn to use the scratchpad for what scratchpads are for — backtracking, double-checking, trying multiple approaches. And the published results back this up. OpenAI’s o1 announcement showed that o1’s performance keeps climbing with more test-time compute; the DeepSeek-R1 paper reports self-correction and verification behaviors emerging from RL alone in their R1-Zero variant — trained with no supervised reasoning traces at all. (The final R1 model adds back a small amount of cold-start data and several SFT/RL stages on top; R1-Zero is the cleaner statement of the emergent-reasoning claim.)
The fix that worked: verifiable rewards
The DeepSeek-R1 recipe is the most public version of this, which is why it’s worth internalizing — it’s the one place you can read what the reward actually was.
Standard RLHF uses a learned reward model that imitates human preferences — useful for “be helpful, be polite,” but a noisy and gameable signal for “is this proof correct.” DeepSeek’s paper describes a rule-based reward system with two parts: an accuracy reward (for math problems, force the model to put its final answer in a known format and check it with a symbolic verifier; for code, run the unit tests) plus a format reward for using the scratchpad correctly. No learned reward model, no human raters in the loop. The community has since latched onto the shorthand reinforcement learning with verifiable rewards (RLVR) for this family of techniques; the term isn’t from the R1 paper itself.
What “emerged” from this training, per the paper, was a set of behaviors nobody hand-wrote: the model started spending more tokens on harder problems, and started stopping mid-solution to reconsider an earlier step — the paper includes a sample where the model interrupts itself to re-examine its own work, which the authors call an “aha moment.” This is the load-bearing claim of the reasoning-model story: long, useful scratchpads are an emergent property of training against a correctness signal, not something you can reliably get from prompting.
The standard account is also that not every domain has a clean verifier. Math and code do. “Write a kind condolence email” doesn’t. This is one reason reasoning models tend to dominate math/code/STEM benchmarks much more than they dominate, say, creative-writing benchmarks. Nobody has published a clean per-domain measurement of how large that gap is, so treat the direction as well-established and the magnitude as task-specific.
The new scaling curve
The conceptual reason this is a Big Deal, not a parlor trick, is the shape of the curve.
Pretraining scaling is the well-known story: more training compute buys you a roughly predictable improvement in loss. Reasoning models opened a second curve: at inference time, holding the model fixed, spending more compute per query also buys measurable improvement on hard tasks — at least up to a point. OpenAI’s original o1 chart showed AIME accuracy climbing roughly linearly against the log of test-time compute.
Two important caveats people skip past:
- Log axes flatter brute force. A linear-log chart can make expensive compute look like clean, predictable scaling. To go from 80% to 90% on a benchmark might cost 10x or 100x more inference compute than 70% to 80%. The line is straight; the underlying resource demand is exponential. Toby Ord’s Inference Scaling and the Log-x Chart is the cleanest argument that these charts are partly a visual rhetorical move.
- Diminishing returns and ceilings exist. Recent work (e.g. arXiv:2502.12215) argues that some o1-like models don’t keep scaling past a budget; they plateau or even degrade if forced to think longer. The “more compute = better” story is real but bounded.
Still, even with those caveats, there’s now a second scaling axis. You can buy quality at training time or at inference time. That changes how labs build models (smaller base + heavier RL run) and how engineers deploy them (route easy queries to a cheap model, hard ones to a reasoning model with a budget).
Where the seams show
A few things worth knowing if you’re shipping with these models:
- You don’t see the raw thinking. Providers deliberately hide the raw scratchpad. OpenAI ships a summarized version in the chat UI and bills the hidden tokens as output tokens via the API. The reasons OpenAI gives include keeping the raw chain unfiltered so it stays monitorable — which means not training policy compliance onto it — and, stated just as plainly, competitive advantage. Either way the practical effect is that you’re paying for tokens you can’t audit.
- Context window pressure. Reasoning tokens consume context window during generation. A reasoning model burning 20k thinking tokens has 20k fewer tokens of room for the rest of the response. Whether that scratchpad survives into the next turn is a provider and model detail rather than a fixed rule — OpenAI now documents ways to persist and reuse reasoning across turns on supported models, so don’t assume it is always thrown away. Within a single turn it competes with your prompt and the visible answer for the same budget, which puts long-context tasks plus heavy reasoning uncomfortably close to the wall.
- Copying the behavior is cheaper than inventing it. A common pattern is to use a frontier reasoning model to generate scratchpads, then fine-tune a smaller base model on those traces. The s1 paper (Muennighoff et al., 2025) reproduced strong test-time-scaling behavior by fine-tuning on just 1,000 curated reasoning traces and adding a simple “budget forcing” trick to control thinking length. Read that as suggestive rather than settled — once a frontier model exists to generate traces, copying its reasoning shape into a smaller model is much cheaper than inventing it from scratch — but that’s interpretation, not the paper’s headline claim. (See the related post on distillation.)
- It’s not magic for everything. Reasoning models buy little or nothing over their non-reasoning siblings on tasks that don’t benefit from deliberation: simple chat, summarization, classification, cheap function calls. Spending 30 seconds and $0.50 to answer “what time is it in Tokyo” is a misuse of the tool.
- The split is probably temporary. It’s not obvious that “reasoning model” will stay a separate SKU long-term. The natural end state is one model with a knob that decides per-query whether to think, and how long. The current split is partly product packaging, partly training-recipe maturity. Which way it settles is an open question, and not one anybody has a track record of calling correctly.
You started with reasoning model = base LLM + RL that rewards correct final answers after a long scratchpad. What did this post add that the
one-liner hides? — + a verifier. Everything else follows from that one
requirement: it’s why math and code led, why the same recipe doesn’t
straightforwardly transfer to “write a kind condolence email,” and why
the answer to the hook is narrower than “models now think.” Your
competition math problem gets a model that keeps working because someone
could mechanically check whether it got the right number.
Check yourself
Before you go — you have a reasoning model and a task with no automatic checker: rewriting support emails to sound warmer. Could you still train a model to spend more test-time compute on it, and what would you have to give up?
Answer
You can spend the compute — nothing stops you from generating a long scratchpad — but you can’t train it the R1 way, because the training signal in that recipe comes from a deterministic checker and there isn’t one here. Your options fall back to a learned reward model (i.e. RLHF), which reintroduces exactly the proxy that verifiable rewards were attractive for avoiding: the model can learn to produce scratchpads that score well rather than emails that are warmer. So the trade is: you keep the extra compute, you give up the clean correctness signal, and you take on reward-hacking risk. That asymmetry is the honest reason reasoning models dominate STEM benchmarks more than creative ones.
And one more: a chart shows accuracy climbing steadily as test-time compute increases, with compute on a log axis. Your product manager reads it as “we can hit 95% by letting it think longer.” What’s the catch, and what would you check first?
Answer
A straight line against log-compute means each equal step up in accuracy costs a multiple of the previous step’s compute — the resource demand is exponential even though the line looks tame. Before promising anything, check two things: where the line actually ends on the published chart (labs plot the range they measured, not an extrapolation), and whether this model plateaus or degrades past a thinking budget, which some o1-style models do. Then price a single query at the budget you’d need. Often the honest answer is “route the hard 5% here and the rest to a cheaper model,” not “turn the dial up.”
Famous related terms
- Chain-of-thought —
CoT = prompt the model to write its reasoning before the answer. The prompt-time ancestor; reasoning models bake the same idea into training. - Test-time compute —
test-time compute = work done per query at inference, not at training. The axis reasoning models scale on. - RLVR (Reinforcement Learning with Verifiable Rewards) —
RLVR = RL + a deterministic checker for the answer instead of a learned reward model. The community’s name for the training trick; the R1 paper describes the recipe (as a “rule-based reward system”) without using the term. - RLHF —
RLHF = SFT + reward model + RL loop. The older post-training recipe RLVR partially replaces for verifiable domains. - Reasoning tokens —
reasoning tokens = output tokens emitted into a hidden scratchpad before the visible answer. What you’re billed for and can’t read. - Distillation —
distillation = teacher model + student trained on teacher's outputs. How reasoning behavior gets copied from a frontier model into a smaller, cheaper one.
Going deeper
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (arXiv:2501.12948) — go here for the only fully public answer to “what exactly do you reward, and what behaviors show up on their own?”
- s1: Simple test-time scaling (arXiv:2501.19393) — read this for the cheap end of the recipe: how far you get by fine-tuning on a small curated set of reasoning traces, and how you control thinking length once you have.
- Revisiting the Test-Time Scaling of o1-like Models (arXiv:2502.12215) — the rabbit hole, for the question the launch charts don’t answer: where does more thinking stop helping?