Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does temperature exist as a knob?

If the model knows the right answer, why is there a dial that asks it to be wrong on purpose?

AI & ML intro Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has only ever seen the word “creativity” next to this knob. The prose below fills in the seams the pictures skip.

1 · The problem

Same prompt, same file, different answer.

“write a birthday message for my sister” not one character changed THE MODEL weights frozen same file both times 1st click regenerate “Happy birthday to the best sister anyone…” “Happy birthday to the most wonderful…” ? Nothing in the model changed. So something after it did.
The weights are frozen — it is the same file on both clicks. Whatever produced the difference sits downstream of the model, and that is the thing temperature is a knob on.

2 · What the model actually hands over

Not a word. A whole landscape of numbers.

Happy birthday to the ___ one number per vocabulary entry, tens of thousands of them 0.31 best 0.22 most … the long tail: every other token in the vocabulary … octopus 0.0000003 The model’s output is this shape. Somebody else has to pick.
At every step the model produces a full probability distribution over its whole vocabulary — not a choice. The step that turns this landscape into one actual word is sampling, and it runs after the model, not inside it.

3 · The obvious rule, and why it fails

“Always take the tallest bar” writes worse text, not better.

always this one every step, every time gives “Happy birthday. I hope you have a great day. I hope you have a great day. I hope you have a great day.” blandest possible sentence — and long generations start looping Maximising likelihood at every step does not produce human-looking text. Holtzman et al. (2019) is the paper that made this concrete — cited in the prose
Taking the maximum at every step is deterministic, which sounds like a feature. It isn’t: the output collapses into safe phrasing and, over a long generation, repeats itself. So samplers pick probabilistically instead — and then the enormous tail becomes the new problem.

4 · The knob

Divide every score by T before turning it into a probability.

p = softmax( logits ÷ T ) the model’s logits never change — only what the picker is handed T = 0.5 0.87 0.12 0.02 sharpened — top token wins almost always T = 1.0 0.66 0.24 0.09 the model’s own distribution, untouched T = 2.0 0.51 0.31 0.19 flattened — the other two are live options now illustrative close-up on three candidates after “Happy birthday to the”, from logits 2.0 / 1.0 / 0.0 — the same numbers worked in the prose
Dividing by a small T magnifies the gaps between scores before they are exponentiated, so the peak pulls further ahead; dividing by a large T shrinks them toward flat. The relative order never changes — only how exaggerated it looks to the picker. T = 0 is the limit: pick the top token, every time.

5 · The seam

Temperature reshapes the range. Its neighbours cut it off.

temperature — reshape every bar still there, heights re-scaled top-p / top-k — truncate deleted before sampling tail removed, survivors renormalised They compose — and providers apply them in different orders. which is why the same 0.7 is not the same intervention on two different APIs
Temperature and the truncation knobs do different jobs, which is why setting both is normal. The order they are applied in is a provider detail that is not always documented — calibrate the number per model rather than carrying it across.

6 · Keep this card

The whole thing on one index card.

temperature = a number that flattens or sharpens the model’s probability landscape + before a token is sampled from it ∴ it never touches the model
Picture to keep: a mountain range of probabilities across the vocabulary, and one dial that pulls the tallest peak up until nothing else is visible or presses the whole range flat — without ever moving the terrain underneath. Your second birthday message differed because the picker rolled a different die, not because the model changed its mind.

Why it exists

You ask a chatbot to write a birthday message for your sister. You don’t love it, so you hit regenerate — same prompt, not one character changed — and a different message comes back. Nothing about the model changed between those two clicks — the weights are frozen; it’s the same file both times. So where did the difference come from? (Hold onto that birthday message; it’s the running example for the rest of the post.)

Somewhere in the API behind that button is a parameter called temperature, on most APIs a float somewhere in the 0-to-2 range, that the docs describe with words like “creativity” or “randomness.” Cranking it up makes the output weirder. Cranking it down makes the output more stable. Neither description tells you what it actually is, and the framing is misleading in a specific way: it suggests the model has a “right answer” and you’re deciding how far to wander from it.

That’s not what’s happening.

A language model doesn’t pick a next token. It produces, on every step, a full probability distribution over its entire vocabulary — tens or hundreds of thousands of numbers, each one a guess at how likely that token is to come next. Mid-way through your birthday message, after Happy birthday to the, the token best might get 0.31, most might get 0.22, octopus might get 0.0000003. The model’s output is that distribution, not a word.

To get text out of that, somebody has to actually pick. That picking step is called sampling, and it’s not part of the model — it’s a separate stage that runs after the forward pass. Temperature is a knob on the sampler, not on the model.

The reason the knob exists is that “always pick the most likely token” — the obvious-seeming default — turns out to produce worse text than picking probabilistically. It loops. It collapses into the same safe phrases. It stops surprising itself, which means it stops surprising you. So real samplers reach for the distribution and reshape it before sampling, and temperature is the simplest control on that reshaping.

Why it matters now

Temperature shows up wherever a model is generating text, which is most places a model is deployed:

The short answer

temperature = a number that flattens or sharpens the model's probability distribution before a token is sampled from it

Picture to keep: a mountain range of probabilities across the vocabulary, and one dial that either pulls the tallest peak up until nothing else is visible, or presses the whole range flat until every hill is a contender — except the dial never moves the terrain. The relative heights the model assigned are fixed; you’re only changing how exaggerated they look to the picker.

At temperature 1, you sample from the model’s distribution as-is. Below 1, you make the peaks taller and the valleys deeper — the most-likely tokens get even more likely, and the long tail gets crushed. Above 1, you flatten the distribution — unlikely tokens get a real shot. Temperature 0 is a limit case: pick the single most-likely token, every time, no randomness left.

How it works

Follow the naive design and watch it break.

Naive attempt: at every step, emit the highest-probability token. It’s the obvious rule, and it’s deterministic, which sounds like a feature. But run it and the birthday message comes out as the blandest possible sentence, and longer generations start looping — the same clause, then the same paragraph, over and over. Holtzman et al. (2019) is the paper that made this concrete: maximizing likelihood at every step does not produce text that looks like the likely text a human would write.

Fix: sample from the distribution instead of taking its maximum. Now best gets picked about 31% of the time and most about 22%, and two regenerations genuinely differ. But raw sampling has the opposite problem — the long tail is enormous, and summed over tens of thousands of vocabulary entries, the junk collectively gets a real share of the draws. Somewhere in a long message, octopus eventually gets its turn.

Fix: reshape the distribution before you sample from it. That’s temperature. And the mechanism is a one-line change inside the softmax that converts the model’s raw output scores — logits — into probabilities.

Vanilla softmax over logits z:

p_i = exp(z_i) / sum_j exp(z_j)

With temperature T, you divide every logit by T first:

p_i = exp(z_i / T) / sum_j exp(z_j / T)

That’s the whole intervention. The model’s logits don’t change. The distribution you sample from does.

Walk through what T does:

A worked example: three candidates after Happy birthday to the, with illustrative logits [2.0, 1.0, 0.0] for best, most, sweetest.

                     best   most  sweetest
T = 1.0:  probs ≈ [  0.66,  0.24,  0.09 ]   moderate preference for the top
T = 0.5:  probs ≈ [  0.87,  0.12,  0.02 ]   strong preference for the top
T = 0.1:  probs ≈ [  1.00,  0.00,  0.00 ]   effectively greedy
T = 2.0:  probs ≈ [  0.51,  0.31,  0.19 ]   much closer to uniform

Same model, same logits — four different sampling regimes. At T = 0.1 every regeneration of the message opens the same way; at T = 2.0 the top choice still wins about half the time, but the other two are now live options instead of rounding errors.

A few things that trip people up:

The deep reason a knob like this has to exist is that there is no single correct distribution to sample from for every task. “Translate this sentence” wants a sharp distribution that always finds the same right answer. “Write a birthday message I haven’t seen before” wants a flat one that lets genuinely different ideas through. The model produces logits; the caller decides which task they’re doing. Temperature is the cheapest control we have for moving between those modes.

You started with temperature = a number that flattens or sharpens the distribution. What did this post add? — + before a token is sampled from it, and that clause is the whole thing: temperature never touches the model. Regenerating gave you a different birthday message because the picker after the model rolled a different die, not because the model changed its mind.

Going deeper

The math here (logit / T inside softmax), the qualitative behavior across the T range, and the fact that “temperature 0” on hosted APIs is near-deterministic but not bit-exact are all well-established. What is not public is the order in which any specific provider composes temperature with top-p / top-k, and whatever internal logit renormalization it does before exposing the knob — providers rarely document either, and both change across model versions. Check the current API reference rather than a blog post.