How does an AI model decide what to say?
It looks like one big choice — you type a question, you get an answer. Underneath it's thousands of tiny choices, made one token at a time, with no plan and no rewind.
On this page
The picture version
The whole idea in six pictures, for a reader who has never wondered what happens between typing a question and watching the answer appear. The prose below fills in the seams the pictures skip.
1 · The problem
Same question, two tabs, two different essays.
2 · The naive way
The obvious story: think first, then type it up.
3 · What one step actually produces
Not a word. A score for every word it knows.
4 · The pick
Something outside the model rolls the dice.
5 · The trick
Every token it picks becomes part of the question.
6 · Keep this card
The whole thing on one index card.
Why it exists
You ask a model “why did the Roman Empire fall?” and watch the answer stream in — word, word, word, at reading speed. It reads like someone who thought about it, picked an angle, and is now typing it up. Ask the identical question in a fresh tab and you get a different essay, sometimes leading with a different cause entirely. That same question is the running example for this post.
Something in there made a decision: which words to use, which facts to assert, which direction to take the response. Where does that decision happen? Is there a moment, somewhere inside the network, where the model “picks the answer”?
You probably assume the model works out the answer and then writes it down. It’s closer to the reverse: the writing is the working out. There is no moment of deciding. The answer assembles itself one token at a time, and nothing is committed until each token is already chosen. By the time the response feels like a thing, several hundred or thousand tiny picks have happened, each conditional on the picks before it, and none of them representing the answer as a whole.
That’s hard to hold onto, because coherent paragraphs look like they were planned. But most of the user-visible quirks of language models fall out of the token-by-token picture once you see it clearly — including why two tabs gave you two different essays.
Why it matters now
Most of the things engineers run into when they ship LLM features stop being mysterious once you stop thinking of generation as a single decision:
- Why the same prompt gives different answers. At the randomized settings most chat products ship with, each token is sampled from a probability distribution. Different rolls, different completions. (Turn sampling off and you mostly stop seeing this — see why temperature zero isn’t deterministic for why “mostly”.)
- Why chain-of-thought prompting helps. The model’s “reasoning” is the trajectory of tokens it writes. Asking it to write the steps changes which token is most likely at each step, and that changes the destination.
- Why prefilling
the assistant’s reply works. Once the model is conditioned on
Sure, here's the JSON: {, the most likely next tokens are JSON, not refusals. You haven’t changed the model — you’ve put it on a different trajectory. - Why the model “can’t take it back” mid-sentence. Tokens already generated are now part of the prompt. There is no rewind. The model is committed to finishing the thing it started.
If your mental model is the model has an answer in mind and is typing it out, the above are confusing. If your mental model is the model is taking one step at a time, with no plan it has access to, they’re inevitable.
The short answer
model decides ≈ forward pass over the whole context + a sampler picks one token + append it + repeat
Picture to keep: someone walking a path with no map, choosing at each fork by rolling weighted dice, where every fork already taken narrows which forks come next. The “answer” is the path, not a destination anyone had in mind. Except — and the last section of this post is about this — “no map” is slightly too strong: the internals appear to carry some anticipation of what’s coming, even though nothing is committed until a token is actually sampled.
There is no other step. Every “decision” you’re seeing is the cumulative effect of doing this hundreds or thousands of times in a row. The forward pass is the model. The sampler is not — it’s a small, cheap function that runs after, and changing it changes everything you see.
How it works
One step
The natural guess is that a forward pass through the network outputs a word. It doesn’t, and that gap is where the whole picture lives.
Start with the prompt. It’s been chopped into tokens — small chunks of text, each with an integer ID. The model takes that whole sequence and runs it through the transformer: dozens of layers of attention and feed-forward computation. The output of all that work, at the very end, is a single vector of logits — one number per token in the model’s vocabulary. The vocabulary is typically 30k–200k tokens.
That vector of logits is what the model “thinks.” It is not an answer. It’s a guess, for this position, at how likely each possible next token is.
Pass the logits through
softmax
to convert them into a probability distribution. On the Roman Empire
question, the first position might come out as The at 0.31, There at
0.22, Rome at 0.11, Historians at 0.04, octopus at 0.0000003, on
down through every entry in the vocabulary. (Those numbers are
illustrative — the point is the shape, not the values.)
This is everything the model has to say about the next token. It is not a token. It is a distribution.
The sampler
So the forward pass can’t finish the job. A distribution isn’t text, and nothing inside the network resolves it. Something outside has to pick. That’s sampling, and it lives outside the model — usually in the inference server. The simplest sampler is “pick the highest-probability token.” The one chat products generally default to is “draw randomly from the distribution after reshaping it” (see temperature and the related top-p / top-k knobs covered in why beam search died for LLMs).
The sampler is the only thing in the entire pipeline that turns the model’s smear of probabilities into a definite token. It’s small, it’s cheap, and its settings are one reason two providers serving the same weights can produce noticeably different outputs — kernel and numerical differences are another.
Repeat
But one token isn’t an answer either. So: once a token is sampled, append it to the sequence. The new sequence — your original prompt plus this one token — goes back through the network. A new forward pass. A new logit vector. A new distribution. A new sample. Append.
Do this until the model emits a special end-of-turn token, or until you hit a length cap.
That is the whole loop. Forward pass, sample, append, forward pass, sample, append. A 500-token answer is 500 of these. A 10,000-token reasoning trace is 10,000. (In practice, the KV cache keeps the key/value vectors already computed for prior tokens, so each new step only has to compute its own — but the logical loop is the same.)
The prompt, chopped into sub-word tokens.
Each position mixes information from earlier ones — thicker line = stronger attention. The right-most output is the vector that predicts the next token.
A score for every token in the vocabulary, softmaxed into probabilities.
Pick one token from the distribution, append it to the context, run the whole loop again.
What “deciding” actually is
But then why is the output coherent at all? If each step is an independent roll, you’d expect the answer to wander off after two sentences and forget it had opened with fiscal collapse. It doesn’t, and the reason is the one structural fact that makes the whole loop work: each forward pass sees everything already chosen. A few consequences are worth pulling out, because they’re the parts that don’t feel right at first:
- The model has no internal “answer” that gets serialized into tokens. The tokens are the computation. Whatever shape the answer eventually has is a property of the trajectory the sampling rolled out, not of some hidden plan. There is no place inside the network where the full answer is sitting, waiting to be written.
- Each step’s distribution depends on every token chosen so far. This is how coherence happens at all. The forward pass at step 437 sees all 436 tokens already committed. It has no choice but to be consistent with them — the patterns the model learned during pretraining strongly weight tokens that “fit” the context.
- Earlier tokens are load-bearing. A token sampled at step 5 reshapes the distribution at step 6, which reshapes the distribution at step 7. Small perturbations early can fork the trajectory hard. This is why temperature 0 looks stable: the same maximum-likelihood path gets followed, so the small perturbations don’t fire.
- There is no rewind. The model can’t un-sample a token. If a wrong fact gets emitted at token 12, by token 50 the model is still continuing from “wrong fact” and will often double down rather than retract — once the prefix exists, the next-token distribution is shaped by it, and locally coherent continuations dominate. This is one mechanical contributor to some hallucinations: not a lie, but a trajectory the sampler stepped onto and that the model is now finishing under continuation pressure.
- Chain-of-thought moves the answer farther down the trajectory. Each intermediate step the model writes becomes part of the context for the next step, so the final answer is conditioned on its own derivation. That prompting for intermediate steps improves accuracy on multi-step reasoning benchmarks, at least in large enough models, is well established (Wei et al., 2022); that the conditioning story above is the whole explanation is my reading of the mechanism, not a settled result.
Where the metaphor breaks
A few honest caveats — places where “no plan at all” is too strong:
- The forward pass isn’t purely about the next token. In at least some cases, internal representations appear to carry constraints about positions several steps ahead. Anthropic’s 2025 interpretability work on Claude 3.5 Haiku (Tracing the thoughts of a large language model) found that when writing rhyming poetry the model appears to commit internally to the line’s end-word before generating the words leading up to it. So a better way to put it is: the model can have anticipations, but those anticipations only become externally committed text when a token is actually sampled. How general this kind of planning is, and at what horizons, is an active research question — I’m describing the shape, not citing a settled result.
- Some inference systems do work several tokens ahead. Speculative decoding drafts several tokens with a small model, then has the big model check them in a single pass, keeping the ones consistent with its own distribution and discarding the rest. This doesn’t change the underlying logic — the big model still has the final say at every position — but mechanically the loop is no longer strictly one-token-at-a-time.
- “Reasoning” models complicate the picture. Several current model families generate a long stretch of intermediate tokens before the visible answer, and bill for them separately; the training recipes behind that behaviour are mostly not public, so I’d be guessing about the details. What’s clear from the outside is that the decision-making is still token-by-token, but the answer you read is now a function of a much longer trajectory you don’t.
- There is no clean theorem for why this loop produces good answers. Empirically it does — see why next-token prediction generalizes — and the intuition is that to predict text well, the model has to model what produced the text. But “good distribution at each step + sampling” being enough for coherent multi-paragraph answers is a thing that works, not a thing that’s been derived from first principles.
The headline still holds: there is no single decision, just many small ones; the model is not executing a plan it has access to, only sampling from a fresh distribution at each position; and the answer you read is the path the sampler walked, not a hidden truth being read out.
You started with `model decides ≈ forward pass + sampler picks one token
- append + repeat
. What did the walk-through add? —+ the append is the whole trick`. Feeding each choice back into the input is what turns a stack of independent guesses into something that sustains one argument about Rome for four paragraphs — and it’s also why your two tabs diverged permanently after one early roll went the other way, and why the model can’t take back a wrong fact it stated at token 12.
Famous related terms
- LLM —
LLM = neural net + "predict the next token" objective at scale. The thing producing the distribution at every step. - Logits —
logits = the raw vector of scores the model emits per vocabulary token. Pre-softmax, pre-sampling. - Softmax —
softmax = exp(logits) / sum(exp(logits)). Turns the raw scores into a probability distribution. - Sampling / temperature —
sampling = turn the distribution into a single token. The step that actually converts model output into text. - Greedy decoding —
greedy = always pick the highest-probability token. The simplest sampler; deterministic on paper, often boring in practice. - Beam search —
beam search = keep the top-k partial sequences alive at each step. Standard in earlier sequence-to-sequence systems like machine translation; largely displaced by plain sampling for open-ended generation. - Chain-of-thought —
CoT = ask the model to write its steps before its answer. Works by changing the trajectory the sampler walks. - Hallucination —
hallucination = a confidently wrong continuation. One mechanical contributor: a trajectory the sampler stepped onto early that the model then has to finish coherently.
Going deeper
- The Curious Case of Neural Text Degeneration (Holtzman et al., 2019) — the primary source for “why isn’t the best sampler just ‘always take the most likely token?’”, answered with the looping, degenerate text that comes out when you do.
- Andrej Karpathy, Let’s build GPT — answers “would this loop still look like a loop if I wrote it myself?”, by building the whole thing from an empty file in about two hours.
- Anthropic, Tracing the thoughts of a large language model (2025) — the rabbit hole, and the source of the rhyme-planning result above: what does the inside of a forward pass look like when a strict next-token-only story stops fitting?