Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does predicting the next token end up doing reasoning?

An LLM is trained on one objective: guess the next token. From that one task, you get translation, code, arithmetic, and arguments. Why is autocomplete this powerful?

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has only ever met autocomplete. The prose below fills in the seams the pictures skip.

1 · The problem

Your phone keyboard plays exactly the same game. It never became this.

your phone, for years now see you to tomorrow nobody has ever called this intelligence a large model, same question def is_prime(n): a working function and this one we argue about = One objective, in both cases: what token comes next? no “understand this” objective, no “reason about this” objective, no “be helpful” objective so the real question is why the keyboard stayed a keyboard
Both systems are trained to do one thing: given some text, guess what follows. The keyboard is the running example for the whole post, because it is the control condition — same game, wildly different outcome.

2 · Why the cheap strategy stalls

Counting what usually follows what runs out almost immediately.

what counting handles “see you to” → “tomorrow” seen a million times; just tally it what counting cannot 7 × 8 =  ? unless it saw that exact string def is_prime(n):  ? the body depends on what primality is is the character from chapter one still alive? nothing in a tally can know Most sentences worth predicting have never been written before. counting harder doesn’t help — the strategy itself is the ceiling
Tallying which words follow which works right up until the text is new, which is most of the time. Scaling up the tally fixes none of these — so whatever the big models are doing, it isn’t a bigger tally.

3 · What the objective actually pushes on

To be less surprised by all of this, you have to model what wrote it.

what it has to keep predicting code one language, then another arithmetic stories that stay consistent chains of consequence forces machinery that captures what produced the text syntax and scope how numbers combine who is in the room and what they want what follows from what The machinery is a byproduct. Nobody asked for it. this is a story that fits the evidence, not a theorem — there is no proof that the objective must produce it
The training signal only ever punishes being surprised by the next token, but the text it is being scored on contains code, arithmetic, argument and narrative. The representations that measurably drive that loss down look like models of what produced the text — an interpretation the field finds compelling, not something anyone has proved.

4 · Where “tasks” come from

Nobody taught it tasks. The prompt just makes the answer the likeliest continuation.

Translate to French: Hello → def is_prime(n): Q: 137 + 248? A: step by step, Bonjour — and the translation is done a primality check — and the coding is done a derivation — and the arithmetic is done Every “task” is a context where finishing the text is the answer. later training makes it treat your message as a request rather than text to continue — but the engine underneath is unchanged
You don’t teach the model a task; you arrange a context in which the likeliest continuation happens to be the thing you wanted. That is why prompting works at all, and why the same machinery does translation, code and arithmetic without any of them being a separate skill.

5 · Why the keyboard stayed a keyboard

Same objective. No room to afford the expensive answer.

a model small enough to run on a phone common phrases rough word order it has to make crude generalisations to fit what it has so it offers you tomorrow, forever a model that fills a data centre it can afford separate machinery for syntax and arithmetic and translation and a thousand more and use each one when it’s relevant The objective never changed. What changed is which answers were affordable.
The representations that predict diverse text well are expensive to learn, and a model has to fit everything it knows into the capacity it has. Treat the sizes as illustrative rather than thresholds anyone measured — the point is that capacity, not objective, is what separates the two.

6 · Keep this card

The whole thing on one index card.

why it generalises = to predict text well you have to model what made it + data diverse enough that shortcuts stop paying + capacity to afford the real answer remove any one of the three and you get a phone keyboard and the limit worth carrying: it is modelling text about the world, not the world — so wherever the writing is consistently wrong, so is the model why it works as well as it does is genuinely not settled: there are scaling-law fits and good post-hoc stories, not a derivation
Picture to keep: a student who must pass an exam covering every subject at once, with only one kind of question — finish this sentence. Cramming phrases earns partial credit; the only way to score well across every subject is to actually learn the subjects. Where it breaks: the student is learning from text about the subjects, so wherever the textbooks are consistently wrong, so is the student.

Why it exists

Your phone keyboard has been finishing your sentences for a decade. You type “see you to” and it offers tomorrow. Nobody has ever called that intelligence, and nobody should. Hold onto that keyboard — it’s the running example for this post, because an LLM is playing the same game. Given a string of tokens, output a probability distribution over what comes next. That’s it. No “understanding” objective. No “reasoning” objective. No “be helpful” objective at the pretraining stage. Just: which token comes next.

And yet what falls out of that one objective is — depending on the day — working code, a passable translation between languages it was never explicitly taught to translate, an argument that holds together for a paragraph, arithmetic on numbers it has never seen in that exact form. Somewhere between “predict the next token” and “write a unit test that passes” there is a gap that, if you’ve never tried to close it, looks absurd.

The interesting question isn’t whether this works. We know it does. The question is why the keyboard stayed a keyboard and the LLM didn’t — same objective, wildly different outcome. And how much of “reasoning” is real, versus us being fooled by fluent text?

Why it matters now

This isn’t philosophy — it’s the thing that decides whether your prompt works. Every time you paste a task into a chatbot and it either nails it or fails in a way that seems arbitrary, the explanation is here. Three concrete places engineers hit this seam without naming it:

If your mental model is “the LLM has been taught to do tasks,” you’ll keep being surprised by what it’s good and bad at. If your mental model is “the LLM has been pressured into a representation of language good enough to predict the next token, and tasks fall out of that,” the surprises mostly stop.

The short answer

next-token prediction generalizes ≈ "to predict text well, you have to model what produced the text"

Picture to keep: a student who must pass an exam covering every subject at once, with only one kind of question — finish this sentence. Cramming phrases gets them partial credit; the only way to score well across every subject is to actually learn the subjects. The picture breaks at one edge: the student is learning from text about the subjects, not the subjects, so wherever the textbooks are consistently wrong, so is the student.

To get good at guessing the next token in arbitrary internet text — with enough capacity to have the option — the model is pushed toward building internal machinery that approximates the things that generated the text: facts, syntax, arithmetic, code semantics, the stance of an author, the structure of an argument. The machinery is the byproduct. Tasks ride on top of it. (Your keyboard is playing the same game without the capacity to take that option, which is the whole difference.)

How it works

The naive way to play this game is what your phone keyboard does: count which words followed which words in a big pile of text, and suggest the most frequent continuation. Cheap, fast, and it genuinely works for “see you to → tomorrow.”

Why it breaks. Push that strategy harder and it hits a wall immediately, because most sentences worth predicting have never appeared before. Counting can’t finish def is_prime(n): correctly, because that function body depends on what primality is, not on which words tend to follow a colon. Counting can’t tell you that 7 × 8 = is followed by 56 rather than 54 unless it happened to see that exact string. Counting has no way to know a character introduced in chapter one is still alive in chapter twelve. Scaling up the counting doesn’t fix any of this; the strategy itself is the limit.

So what does fix it? Start with what the loss is actually measuring. During pretraining, the model is shown enormous amounts of text and asked, for every position, “what’s the next token?” The training signal — cross-entropy loss — punishes it in proportion to how surprised it was by the right token. Lower loss means it was less surprised, on average, across everything in the training set.

Now think about what “everything in the training set” contains. To get loss materially down across all of it — not just better than chance, which counting already achieves — the model has to genuinely improve on:

It’s hard to see a shortcut for any of these that doesn’t, at some level, model the thing being described. A model that has memorized text but has no notion of arithmetic can’t reliably continue novel arithmetic. A model that has no notion of variable scope can’t reliably continue novel code. The objective is “predict the next token,” but the representations that appear to drive the loss down across that whole corpus are ones that, in some compressed form, capture what produced the text. This framing — prediction is compression, compression requires modeling — is one Ilya Sutskever has argued across several talks and interviews rather than in a canonical write-up, so take the attribution loosely. It’s intuition, not a proof, but it matches what we see: the more diverse and structured the data, the more structure the model is forced to internalize to keep predicting well.

Tasks as conditional continuations

Once you have a model that’s good at next-token prediction, you don’t “teach it tasks” — you arrange the prompt so the right continuation is the task’s answer. The prompt sets a context in which the most likely continuation, according to the patterns the model learned, is the thing you wanted.

This is why in-context learning works at all. The model isn’t learning a task in any usual sense; the prompt is steering an already-built distribution toward the slice that produces the right kind of continuation. RLHF and instruction tuning then re-shape that distribution further so the model treats user messages as task specs, but the engine underneath is still next-token prediction.

Why scale is doing the heavy lifting

Here’s where the keyboard finally parts company with the LLM. Your phone’s model is small — it has to run on a phone, offline, in milliseconds — and a small model trained on the same objective doesn’t get you working code. The reason — best as anyone can tell — is that the representations needed to predict diverse text well are expensive to learn. Take the scales as illustrative, not as thresholds anyone measured: a model with hundreds of millions of parameters has to make crude generalizations to fit its capacity, while one with hundreds of billions can afford features for syntax and arithmetic and translation and a thousand other regularities in the data, and use each one when relevant.

This is the rough shape of scaling laws: loss falls smoothly with more compute and data. But specific capabilities — arithmetic past two digits, multi-step reasoning, following instructions — appear to switch on more abruptly at certain scales. Whether that abruptness is real or partly a measurement artifact (some of the “emergence” results have been re-analyzed and softened) is still actively debated. The honest version: average loss goes down smoothly, and a lot of capabilities ride on that, but the exact mapping from “loss” to “capability” is messier than the early emergence narrative suggested.

Where the story breaks down

The “to predict text you must model the world” framing is the right intuition, but taken too far it becomes wrong in load-bearing ways:

The headline still holds: a single, almost embarrassingly simple objective, applied to enough text with enough capacity, ends up forcing the model to assemble most of what we’d recognize as linguistic and semi-conceptual structure. Tasks are then prompts that sample from that structure. Your keyboard plays the same game with a model too small to be forced into any of it — which is why it will offer you tomorrow forever and never offer you a working function.

You started with next-token prediction generalizes ≈ "to predict text well, you have to model what produced the text". What did this post add that the line leaves out? — + only under pressure. The objective alone doesn’t force anything; it’s the objective plus data diverse enough that shortcuts stop paying plus capacity to afford the real representations. Remove any one and you get a phone keyboard. That’s also the honest limit of the claim: it’s a story that fits the evidence, not a theorem that predicted it.

Check yourself

Before you go — a friend argues that because LLMs are “just predicting the next token,” they can’t be doing anything that deserves the word reasoning. Where does that argument go wrong, and where is it right?

Answer

It goes wrong in treating the objective as a ceiling on the mechanism. “Predict the next token” describes what the model is scored on, not what it had to build internally to score well. Since no counting-based shortcut can continue novel arithmetic or novel code, whatever machinery does drive the loss down has to approximate the things that generated the text. Calling that “just prediction” is like calling a chess engine “just picking a legal move.”

It’s right in two places worth conceding. First, the model is modeling text about the world, so it inherits the internet’s consistent errors. Second, some behavior that looks like reasoning is pattern matching that coincides with reasoning on the training distribution — which is why perturbing surface form (renaming variables, changing the numbers) sometimes collapses performance. The honest position is that the objective doesn’t rule reasoning out, and fluent output doesn’t demonstrate it.

And one more — you fine-tune a small model on nothing but chess games in algebraic notation, and it starts making legal moves it never saw. Does this post predict that?

Answer

Yes, and it’s a clean instance of the mechanism on a narrow corpus. Legal chess notation is highly structured — the set of valid next moves depends on the board state, which depends on every move so far. There is no counting shortcut that gets you there, because the space of positions is far larger than any training set. So the only way to lower the loss is to internalize something that tracks board state. Note what the post also predicts, though: this model will be good at chess notation and nothing else. Diversity of data is what buys breadth; this one bought depth in a single narrow structure. That’s the same trade in miniature.

Going deeper