Why does predicting the next token end up doing reasoning?
An LLM is trained on one objective: guess the next token. From that one task, you get translation, code, arithmetic, and arguments. Why is autocomplete this powerful?
On this page
The picture version
Six pictures for a reader who has only ever met autocomplete. The prose below fills in the seams the pictures skip.
1 · The problem
Your phone keyboard plays exactly the same game. It never became this.
2 · Why the cheap strategy stalls
Counting what usually follows what runs out almost immediately.
3 · What the objective actually pushes on
To be less surprised by all of this, you have to model what wrote it.
4 · Where “tasks” come from
Nobody taught it tasks. The prompt just makes the answer the likeliest continuation.
5 · Why the keyboard stayed a keyboard
Same objective. No room to afford the expensive answer.
6 · Keep this card
The whole thing on one index card.
Why it exists
Your phone keyboard has been finishing your sentences for a decade. You type “see you to” and it offers tomorrow. Nobody has ever called that intelligence, and nobody should. Hold onto that keyboard — it’s the running example for this post, because an LLM is playing the same game. Given a string of tokens, output a probability distribution over what comes next. That’s it. No “understanding” objective. No “reasoning” objective. No “be helpful” objective at the pretraining stage. Just: which token comes next.
And yet what falls out of that one objective is — depending on the day — working code, a passable translation between languages it was never explicitly taught to translate, an argument that holds together for a paragraph, arithmetic on numbers it has never seen in that exact form. Somewhere between “predict the next token” and “write a unit test that passes” there is a gap that, if you’ve never tried to close it, looks absurd.
The interesting question isn’t whether this works. We know it does. The question is why the keyboard stayed a keyboard and the LLM didn’t — same objective, wildly different outcome. And how much of “reasoning” is real, versus us being fooled by fluent text?
Why it matters now
This isn’t philosophy — it’s the thing that decides whether your prompt works. Every time you paste a task into a chatbot and it either nails it or fails in a way that seems arbitrary, the explanation is here. Three concrete places engineers hit this seam without naming it:
- Why does chain-of-thought prompting help at all? If the model “knows” the answer, why does asking for the steps change it? The standard explanation — a good story that fits the evidence rather than a proven mechanism — is that predicting “the answer is 47” in one shot is a different computation than predicting it after writing out the reasoning trace, because the intermediate tokens are extra state the model can condition on.
- Why are LLMs strong at things their trainers didn’t aim at, and brittle at things that look easier? Counting characters in a word is hard for them; explaining a 19th-century court ruling is easy. Both of these stop being mysterious once you remember what task the model was actually trained on.
- Why does emergence happen at scale? The objective doesn’t change as you add parameters and data. What changes is which patterns are cheap enough for the model to learn under that single objective.
If your mental model is “the LLM has been taught to do tasks,” you’ll keep being surprised by what it’s good and bad at. If your mental model is “the LLM has been pressured into a representation of language good enough to predict the next token, and tasks fall out of that,” the surprises mostly stop.
The short answer
next-token prediction generalizes ≈ "to predict text well, you have to model what produced the text"
Picture to keep: a student who must pass an exam covering every subject at once, with only one kind of question — finish this sentence. Cramming phrases gets them partial credit; the only way to score well across every subject is to actually learn the subjects. The picture breaks at one edge: the student is learning from text about the subjects, not the subjects, so wherever the textbooks are consistently wrong, so is the student.
To get good at guessing the next token in arbitrary internet text — with enough capacity to have the option — the model is pushed toward building internal machinery that approximates the things that generated the text: facts, syntax, arithmetic, code semantics, the stance of an author, the structure of an argument. The machinery is the byproduct. Tasks ride on top of it. (Your keyboard is playing the same game without the capacity to take that option, which is the whole difference.)
How it works
The naive way to play this game is what your phone keyboard does: count which words followed which words in a big pile of text, and suggest the most frequent continuation. Cheap, fast, and it genuinely works for “see you to → tomorrow.”
Why it breaks. Push that strategy harder and it hits a wall
immediately, because most sentences worth predicting have never
appeared before. Counting can’t finish def is_prime(n): correctly,
because that function body depends on what primality is, not on which
words tend to follow a colon. Counting can’t tell you that 7 × 8 = is
followed by 56 rather than 54 unless it happened to see that exact
string. Counting has no way to know a character introduced in chapter
one is still alive in chapter twelve. Scaling up the counting doesn’t
fix any of this; the strategy itself is the limit.
So what does fix it? Start with what the loss is actually measuring. During pretraining, the model is shown enormous amounts of text and asked, for every position, “what’s the next token?” The training signal — cross-entropy loss — punishes it in proportion to how surprised it was by the right token. Lower loss means it was less surprised, on average, across everything in the training set.
Now think about what “everything in the training set” contains. To get loss materially down across all of it — not just better than chance, which counting already achieves — the model has to genuinely improve on:
- Code, where the next token depends on syntax, scope, and what the function is supposed to do.
- Translations, where one language’s sentence is followed by another’s.
- Arithmetic strings, where
7 × 8 =is followed by56more often than by54. - Stories, where character names persist and Chekhov’s gun goes off in Act III.
- Reasoning chains, where each line is the consequence of the last.
It’s hard to see a shortcut for any of these that doesn’t, at some level, model the thing being described. A model that has memorized text but has no notion of arithmetic can’t reliably continue novel arithmetic. A model that has no notion of variable scope can’t reliably continue novel code. The objective is “predict the next token,” but the representations that appear to drive the loss down across that whole corpus are ones that, in some compressed form, capture what produced the text. This framing — prediction is compression, compression requires modeling — is one Ilya Sutskever has argued across several talks and interviews rather than in a canonical write-up, so take the attribution loosely. It’s intuition, not a proof, but it matches what we see: the more diverse and structured the data, the more structure the model is forced to internalize to keep predicting well.
Tasks as conditional continuations
Once you have a model that’s good at next-token prediction, you don’t “teach it tasks” — you arrange the prompt so the right continuation is the task’s answer. The prompt sets a context in which the most likely continuation, according to the patterns the model learned, is the thing you wanted.
- “Translate to French: Hello → ” biases the next tokens toward Bonjour. Not because that exact string appeared and continued that way — it’s a learned mapping between the languages, which is why it also works on sentences nobody has ever written.
- “def is_prime(n):” biases the next tokens toward a Python primality check.
- “Q: What’s 137 + 248? A: Let’s think step by step.” biases toward a step-by-step derivation, which on many tasks lands on the right answer more often than a one-shot guess. (Not all of them — chain-of-thought is a reliable tendency, not a guarantee, and there are tasks where it makes things worse.)
This is why in-context learning works at all. The model isn’t learning a task in any usual sense; the prompt is steering an already-built distribution toward the slice that produces the right kind of continuation. RLHF and instruction tuning then re-shape that distribution further so the model treats user messages as task specs, but the engine underneath is still next-token prediction.
Why scale is doing the heavy lifting
Here’s where the keyboard finally parts company with the LLM. Your phone’s model is small — it has to run on a phone, offline, in milliseconds — and a small model trained on the same objective doesn’t get you working code. The reason — best as anyone can tell — is that the representations needed to predict diverse text well are expensive to learn. Take the scales as illustrative, not as thresholds anyone measured: a model with hundreds of millions of parameters has to make crude generalizations to fit its capacity, while one with hundreds of billions can afford features for syntax and arithmetic and translation and a thousand other regularities in the data, and use each one when relevant.
This is the rough shape of scaling laws: loss falls smoothly with more compute and data. But specific capabilities — arithmetic past two digits, multi-step reasoning, following instructions — appear to switch on more abruptly at certain scales. Whether that abruptness is real or partly a measurement artifact (some of the “emergence” results have been re-analyzed and softened) is still actively debated. The honest version: average loss goes down smoothly, and a lot of capabilities ride on that, but the exact mapping from “loss” to “capability” is messier than the early emergence narrative suggested.
Where the story breaks down
The “to predict text you must model the world” framing is the right intuition, but taken too far it becomes wrong in load-bearing ways:
- The model is modeling text about the world, not the world. If the internet is wrong about something in a consistent way, the model will be wrong with it. Hallucinated citations are a classic case: fluent text where citations go is a strong pattern; real citations are a weaker one.
- Some patterns that look like reasoning are pattern matching that happens to coincide with reasoning on the training distribution and comes apart off it. This is why benchmarks that perturb surface form (rename variables, change numbers) sometimes drop scores sharply.
- Tokenization leaks through. In the subword-tokenized models that dominate production, counting characters in a word is hard because the model doesn’t receive characters; it receives tokens. (Byte- and character-level architectures exist and don’t have this problem; they’re rare in production.) No amount of next-token training fixes a representation that hides the unit you’re being asked about.
- “Why does this work as well as it does” is genuinely not fully understood. There’s no clean theorem saying next-token prediction on internet-scale text must yield code-writing assistants. We have scaling-law fits and post-hoc stories. The post-hoc stories are good. They are not a derivation.
The headline still holds: a single, almost embarrassingly simple objective, applied to enough text with enough capacity, ends up forcing the model to assemble most of what we’d recognize as linguistic and semi-conceptual structure. Tasks are then prompts that sample from that structure. Your keyboard plays the same game with a model too small to be forced into any of it — which is why it will offer you tomorrow forever and never offer you a working function.
You started with next-token prediction generalizes ≈ "to predict text well, you have to model what produced the text". What did this post
add that the line leaves out? — + only under pressure. The objective
alone doesn’t force anything; it’s the objective plus data diverse
enough that shortcuts stop paying plus capacity to afford the real
representations. Remove any one and you get a phone keyboard. That’s
also the honest limit of the claim: it’s a story that fits the
evidence, not a theorem that predicted it.
Check yourself
Before you go — a friend argues that because LLMs are “just predicting the next token,” they can’t be doing anything that deserves the word reasoning. Where does that argument go wrong, and where is it right?
Answer
It goes wrong in treating the objective as a ceiling on the mechanism. “Predict the next token” describes what the model is scored on, not what it had to build internally to score well. Since no counting-based shortcut can continue novel arithmetic or novel code, whatever machinery does drive the loss down has to approximate the things that generated the text. Calling that “just prediction” is like calling a chess engine “just picking a legal move.”
It’s right in two places worth conceding. First, the model is modeling text about the world, so it inherits the internet’s consistent errors. Second, some behavior that looks like reasoning is pattern matching that coincides with reasoning on the training distribution — which is why perturbing surface form (renaming variables, changing the numbers) sometimes collapses performance. The honest position is that the objective doesn’t rule reasoning out, and fluent output doesn’t demonstrate it.
And one more — you fine-tune a small model on nothing but chess games in algebraic notation, and it starts making legal moves it never saw. Does this post predict that?
Answer
Yes, and it’s a clean instance of the mechanism on a narrow corpus. Legal chess notation is highly structured — the set of valid next moves depends on the board state, which depends on every move so far. There is no counting shortcut that gets you there, because the space of positions is far larger than any training set. So the only way to lower the loss is to internalize something that tracks board state. Note what the post also predicts, though: this model will be good at chess notation and nothing else. Diversity of data is what buys breadth; this one bought depth in a single narrow structure. That’s the same trade in miniature.
Famous related terms
- LLM —
LLM = neural net + "predict the next token" objective at scale. The thing this whole post is unpacking. - In-context learning —
in-context learning = prompt + frozen weights + a continuation that happens to be the task answer. The mechanism by which “tasks” exist at all. - Chain-of-thought —
CoT = prompt the model to write its steps before its answer. Works because next-token prediction over a derivation is a different (often better-behaved) computation than one-shot answer prediction. - Scaling laws —
scaling laws ≈ loss falls predictably as you add parameters, data, and compute. The empirical reason “more of the same objective” keeps producing better models. - Emergence —
emergence = capability that appears sharply at scale. A real-feeling pattern, with caveats; the average loss curve is smoother than the capability curves it carries. - Pretraining vs fine-tuning —
pretraining = build representations from scratch;fine-tuning = rent existing ones for a task. Pretraining buys the representations; fine-tuning rents them for a specific task.
Going deeper
- Kaplan et al., “Scaling Laws for Neural Language Models” (2020) — the primary source for “how predictable is ‘more of the same objective,’ actually?” Pair it with Hoffmann et al., “Training Compute-Optimal Large Language Models” (Chinchilla, 2022), which answers the follow-up: more of what, parameters or data?
- Schaeffer, Miranda, Koyejo, “Are Emergent Abilities of Large Language Models a Mirage?” (NeurIPS 2023) — the explainer for “how much of ‘emergence’ is the model and how much is the metric?” Read it alongside the original emergence papers, not instead of them.
- Ilya Sutskever’s talks on prediction-as-compression — the rabbit hole for the strong form of the “to predict well you must model” intuition. A gap worth naming: there is no canonical talk to point at and no rigorous write-up of the argument; it lives across several interviews and lectures as intuition, not proof.