Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

How does an AI model decide what to say?

It looks like one big choice — you type a question, you get an answer. Underneath it's thousands of tiny choices, made one token at a time, with no plan and no rewind.

AI & ML intro Apr 30, 2026 · updated Aug 24, 2026 · 13 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never wondered what happens between typing a question and watching the answer appear. The prose below fills in the seams the pictures skip.

1 · The problem

Same question, two tabs, two different essays.

“why did the Roman Empire fall?” asked once same model, same settings asked again tab 1 tab 2 opens with fiscal collapse opens with a different cause ? where did the two answers split?
Nothing was different about the model, the prompt or the settings — and yet the two answers part company somewhere. That somewhere is what this post is about.

2 · The naive way

The obvious story: think first, then type it up.

1. work out the answer somewhere inside the network then 2. type it out word by word, for you THERE IS NO STEP 1
It is closer to the reverse: the writing is the working out. There is no moment where the model has the finished answer in hand and starts transcribing it.

3 · What one step actually produces

Not a word. A score for every word it knows.

the model one pass over everything so far outputs The There Rome Historians octopus still on the list, just tiny one number per token in the vocabulary: 30k–200k of them This is not a word. It is a whole distribution.
A single pass through the network never emits text. It emits how likely every possible next token is, right here, right now. Bar lengths are illustrative.

4 · The pick

Something outside the model rolls the dice.

inside the model The There Rome … handed to outside the model THE SAMPLER rolls weighted dice one token The chosen Nothing inside the network resolves the distribution.
The sampler is the only step that turns a smear of probabilities into one definite token — and it is not part of the model. Change it and everything you read changes.

5 · The trick

Every token it picks becomes part of the question.

un-sampling a token — never happens why did the Roman Empire fall? The Roman Empire your prompt already generated — now part of the input the whole context goes in each time run the whole thing again over every token above one more token, appended to the right-hand end The answer so far is now part of the question.
The append is the whole trick. Feeding each choice back in is what keeps four paragraphs on one argument — and why a wrong fact at token 12 cannot be taken back at token 50.

6 · Keep this card

The whole thing on one index card.

the model decides = one pass over everything so far + a dice roll to pick one token + paste it on the end + do it all again
Picture to keep: someone walking a path with no map, choosing at each fork by rolling weighted dice — where every fork already taken narrows which forks come next. The answer is the path, not a destination anyone had in mind.

Why it exists

You ask a model “why did the Roman Empire fall?” and watch the answer stream in — word, word, word, at reading speed. It reads like someone who thought about it, picked an angle, and is now typing it up. Ask the identical question in a fresh tab and you get a different essay, sometimes leading with a different cause entirely. That same question is the running example for this post.

Something in there made a decision: which words to use, which facts to assert, which direction to take the response. Where does that decision happen? Is there a moment, somewhere inside the network, where the model “picks the answer”?

You probably assume the model works out the answer and then writes it down. It’s closer to the reverse: the writing is the working out. There is no moment of deciding. The answer assembles itself one token at a time, and nothing is committed until each token is already chosen. By the time the response feels like a thing, several hundred or thousand tiny picks have happened, each conditional on the picks before it, and none of them representing the answer as a whole.

That’s hard to hold onto, because coherent paragraphs look like they were planned. But most of the user-visible quirks of language models fall out of the token-by-token picture once you see it clearly — including why two tabs gave you two different essays.

Why it matters now

Most of the things engineers run into when they ship LLM features stop being mysterious once you stop thinking of generation as a single decision:

If your mental model is the model has an answer in mind and is typing it out, the above are confusing. If your mental model is the model is taking one step at a time, with no plan it has access to, they’re inevitable.

The short answer

model decides ≈ forward pass over the whole context + a sampler picks one token + append it + repeat

Picture to keep: someone walking a path with no map, choosing at each fork by rolling weighted dice, where every fork already taken narrows which forks come next. The “answer” is the path, not a destination anyone had in mind. Except — and the last section of this post is about this — “no map” is slightly too strong: the internals appear to carry some anticipation of what’s coming, even though nothing is committed until a token is actually sampled.

There is no other step. Every “decision” you’re seeing is the cumulative effect of doing this hundreds or thousands of times in a row. The forward pass is the model. The sampler is not — it’s a small, cheap function that runs after, and changing it changes everything you see.

How it works

One step

The natural guess is that a forward pass through the network outputs a word. It doesn’t, and that gap is where the whole picture lives.

Start with the prompt. It’s been chopped into tokens — small chunks of text, each with an integer ID. The model takes that whole sequence and runs it through the transformer: dozens of layers of attention and feed-forward computation. The output of all that work, at the very end, is a single vector of logits — one number per token in the model’s vocabulary. The vocabulary is typically 30k–200k tokens.

That vector of logits is what the model “thinks.” It is not an answer. It’s a guess, for this position, at how likely each possible next token is.

Pass the logits through softmax to convert them into a probability distribution. On the Roman Empire question, the first position might come out as The at 0.31, There at 0.22, Rome at 0.11, Historians at 0.04, octopus at 0.0000003, on down through every entry in the vocabulary. (Those numbers are illustrative — the point is the shape, not the values.)

This is everything the model has to say about the next token. It is not a token. It is a distribution.

The sampler

So the forward pass can’t finish the job. A distribution isn’t text, and nothing inside the network resolves it. Something outside has to pick. That’s sampling, and it lives outside the model — usually in the inference server. The simplest sampler is “pick the highest-probability token.” The one chat products generally default to is “draw randomly from the distribution after reshaping it” (see temperature and the related top-p / top-k knobs covered in why beam search died for LLMs).

The sampler is the only thing in the entire pipeline that turns the model’s smear of probabilities into a definite token. It’s small, it’s cheap, and its settings are one reason two providers serving the same weights can produce noticeably different outputs — kernel and numerical differences are another.

Repeat

But one token isn’t an answer either. So: once a token is sampled, append it to the sequence. The new sequence — your original prompt plus this one token — goes back through the network. A new forward pass. A new logit vector. A new distribution. A new sample. Append.

Do this until the model emits a special end-of-turn token, or until you hit a length cap.

That is the whole loop. Forward pass, sample, append, forward pass, sample, append. A 500-token answer is 500 of these. A 10,000-token reasoning trace is 10,000. (In practice, the KV cache keeps the key/value vectors already computed for prior tokens, so each new step only has to compute its own — but the logical loop is the same.)

Input tokens

The prompt, chopped into sub-word tokens.

The cat sat on the mat
Embed

Each token ID is looked up in a learned table; the result is its vector — a position in “meaning” space.

The
cat
sat
on
the
Attention + transformer layers

Each position mixes information from earlier ones — thicker line = stronger attention. The right-most output is the vector that predicts the next token.

Next-token distribution

A score for every token in the vocabulary, softmaxed into probabilities.

mat 0.42
floor 0.21
rug 0.12
bed 0.09
… …
Sample → append → repeat

Pick one token from the distribution, append it to the context, run the whole loop again.

mat ↑ appended to step 1; the loop runs again
The same loop, on a shorter sentence than the running example so the vocabulary fits on screen: each token is embedded into a vector, attention mixes information across positions, the head emits a distribution over the vocabulary, the sampler picks one token, it’s appended to the context, and the loop runs again. Numbers and shapes are illustrative.

What “deciding” actually is

But then why is the output coherent at all? If each step is an independent roll, you’d expect the answer to wander off after two sentences and forget it had opened with fiscal collapse. It doesn’t, and the reason is the one structural fact that makes the whole loop work: each forward pass sees everything already chosen. A few consequences are worth pulling out, because they’re the parts that don’t feel right at first:

Where the metaphor breaks

A few honest caveats — places where “no plan at all” is too strong:

The headline still holds: there is no single decision, just many small ones; the model is not executing a plan it has access to, only sampling from a fresh distribution at each position; and the answer you read is the path the sampler walked, not a hidden truth being read out.

You started with `model decides ≈ forward pass + sampler picks one token

Going deeper