Why does temperature exist as a knob?
If the model knows the right answer, why is there a dial that asks it to be wrong on purpose?
On this page
The picture version
Six pictures for a reader who has only ever seen the word “creativity” next to this knob. The prose below fills in the seams the pictures skip.
1 · The problem
Same prompt, same file, different answer.
2 · What the model actually hands over
Not a word. A whole landscape of numbers.
3 · The obvious rule, and why it fails
“Always take the tallest bar” writes worse text, not better.
4 · The knob
Divide every score by T before turning it into a probability.
5 · The seam
Temperature reshapes the range. Its neighbours cut it off.
6 · Keep this card
The whole thing on one index card.
Why it exists
You ask a chatbot to write a birthday message for your sister. You don’t love it, so you hit regenerate — same prompt, not one character changed — and a different message comes back. Nothing about the model changed between those two clicks — the weights are frozen; it’s the same file both times. So where did the difference come from? (Hold onto that birthday message; it’s the running example for the rest of the post.)
Somewhere in the API behind that button is a parameter called temperature,
on most APIs a float somewhere in the 0-to-2 range, that the docs
describe with words like
“creativity” or “randomness.” Cranking it up makes the output weirder.
Cranking it down makes the output more stable. Neither description tells you
what it actually is, and the framing is misleading in a specific way: it
suggests the model has a “right answer” and you’re deciding how far to
wander from it.
That’s not what’s happening.
A language model doesn’t pick a next token. It produces, on every step, a
full probability distribution over its entire vocabulary — tens or hundreds
of thousands of numbers, each one a guess at how likely that token is to
come next. Mid-way through your birthday message, after Happy birthday to the, the token best might get 0.31, most might get 0.22, octopus
might get 0.0000003. The model’s output is that distribution, not a word.
To get text out of that, somebody has to actually pick. That picking step is called sampling, and it’s not part of the model — it’s a separate stage that runs after the forward pass. Temperature is a knob on the sampler, not on the model.
The reason the knob exists is that “always pick the most likely token” — the obvious-seeming default — turns out to produce worse text than picking probabilistically. It loops. It collapses into the same safe phrases. It stops surprising itself, which means it stops surprising you. So real samplers reach for the distribution and reshape it before sampling, and temperature is the simplest control on that reshaping.
Why it matters now
Temperature shows up wherever a model is generating text, which is most places a model is deployed:
- Determinism vs. variety in agents. A coding agent that re-runs the same prompt and gets a different patch every time is hard to debug. Set temperature to 0 and (mostly) the same prompt yields the same output. That’s not a personality choice; it’s an engineering one.
- Evals are temperature-sensitive. A benchmark score reported “at temperature 0.7” and one at “temperature 0” can differ noticeably for the same model. If a leaderboard doesn’t tell you, you’re comparing apples to oranges.
- Creative tasks need spread. Brainstorming, fiction, marketing copy — ask for ten variants at temperature 0 and you get one variant ten times, because the model keeps reaching for the single most-likely continuation. You want the distribution to actually be a distribution.
- “Hallucination” interacts with temperature, but isn’t caused by it. Higher temperature widens the set of tokens the model is willing to emit, which can let through factually wrong ones — but a confident wrong answer at temperature 0 is just as much a hallucination. Temperature isn’t a truthfulness knob.
The short answer
temperature = a number that flattens or sharpens the model's probability distribution before a token is sampled from it
Picture to keep: a mountain range of probabilities across the vocabulary, and one dial that either pulls the tallest peak up until nothing else is visible, or presses the whole range flat until every hill is a contender — except the dial never moves the terrain. The relative heights the model assigned are fixed; you’re only changing how exaggerated they look to the picker.
At temperature 1, you sample from the model’s distribution as-is. Below 1, you make the peaks taller and the valleys deeper — the most-likely tokens get even more likely, and the long tail gets crushed. Above 1, you flatten the distribution — unlikely tokens get a real shot. Temperature 0 is a limit case: pick the single most-likely token, every time, no randomness left.
How it works
Follow the naive design and watch it break.
Naive attempt: at every step, emit the highest-probability token. It’s the obvious rule, and it’s deterministic, which sounds like a feature. But run it and the birthday message comes out as the blandest possible sentence, and longer generations start looping — the same clause, then the same paragraph, over and over. Holtzman et al. (2019) is the paper that made this concrete: maximizing likelihood at every step does not produce text that looks like the likely text a human would write.
Fix: sample from the distribution instead of taking its maximum. Now
best gets picked about 31% of the time and most about 22%, and two
regenerations genuinely differ. But raw sampling has the opposite problem —
the long tail is enormous, and summed over tens of thousands of vocabulary
entries, the junk collectively gets a real share of the draws. Somewhere in a
long message, octopus eventually gets its turn.
Fix: reshape the distribution before you sample from it. That’s temperature. And the mechanism is a one-line change inside the softmax that converts the model’s raw output scores — logits — into probabilities.
Vanilla softmax over logits z:
p_i = exp(z_i) / sum_j exp(z_j)
With temperature T, you divide every logit by T first:
p_i = exp(z_i / T) / sum_j exp(z_j / T)
That’s the whole intervention. The model’s logits don’t change. The distribution you sample from does.
Walk through what T does:
T = 1— divide-by-1 is a no-op. You sample from the model’s native distribution.T < 1(e.g. 0.2) — dividing by a small number magnifies the gaps between logits before they get exponentiated. The biggest logit pulls even further ahead. The distribution becomes spiky. Sampling from a spiky distribution almost always gives you the top token.T → 0— the limit of the above. Assuming one token is strictly ahead (exact ties need a tie-break rule), it ends up with probability 1 and everything else with 0. This is mathematically argmax, often called “greedy decoding.”T > 1(e.g. 1.5) — dividing by a number bigger than 1 shrinks the gaps between logits. The distribution flattens toward uniform. Rare tokens become plausible. At very highT, output approaches gibberish because nearly every token is roughly equally likely.
A worked example: three candidates after Happy birthday to the, with
illustrative logits [2.0, 1.0, 0.0] for best, most, sweetest.
best most sweetest
T = 1.0: probs ≈ [ 0.66, 0.24, 0.09 ] moderate preference for the top
T = 0.5: probs ≈ [ 0.87, 0.12, 0.02 ] strong preference for the top
T = 0.1: probs ≈ [ 1.00, 0.00, 0.00 ] effectively greedy
T = 2.0: probs ≈ [ 0.51, 0.31, 0.19 ] much closer to uniform
Same model, same logits — four different sampling regimes. At T = 0.1 every
regeneration of the message opens the same way; at T = 2.0 the top choice
still wins about half the time, but the other two are now live options
instead of rounding errors.
A few things that trip people up:
- Temperature 0 is not always actually deterministic. On the math, yes: argmax is a function. In practice, hosted APIs run on GPUs in batched, parallel kernels where floating-point summations are not order-stable, and near-ties in the top logits can resolve differently across runs. So “temperature 0” usually means “nearly deterministic” in production — the longer version is its own post. The user-visible takeaway: don’t bet correctness on bit-exact reproducibility.
- Temperature is not the only sampler knob. Most APIs also expose one or both of top-p (nucleus) and top-k. These chop off the tail of the distribution before sampling. Temperature reshapes the distribution; top-p/top-k truncate it. They compose, and the order matters. Different providers apply them in different orders, and the docs don’t always say which.
- Different APIs use different scales. The accepted range differs by
provider and sometimes between two APIs from the same provider — check the
reference for the endpoint you’re actually calling rather than assuming.
Either way, the same numeric value (
0.7) is not the same intervention across providers, because their underlying logit distributions and any internal renormalization differ. Calibrate per model. - Temperature does nothing during training. It’s purely an inference- time sampling parameter. The model is the same model whether you set temperature to 0 or 2; only the picking step changes.
- It’s not really “creativity.” That word makes it sound like a deeper cognitive setting. It isn’t. It’s a slider on how much you trust the model’s top guess versus how much you let the long tail in.
The deep reason a knob like this has to exist is that there is no single correct distribution to sample from for every task. “Translate this sentence” wants a sharp distribution that always finds the same right answer. “Write a birthday message I haven’t seen before” wants a flat one that lets genuinely different ideas through. The model produces logits; the caller decides which task they’re doing. Temperature is the cheapest control we have for moving between those modes.
You started with temperature = a number that flattens or sharpens the distribution. What did this post add? — + before a token is sampled from it, and that clause is the whole thing: temperature never touches the
model. Regenerating gave you a different birthday message because the picker
after the model rolled a different die, not because the model changed its
mind.
Famous related terms
- Softmax —
softmax = exp(logits) / sum(exp(logits))— the function temperature is parameterizing. Turns arbitrary scores into a probability distribution. - Greedy decoding —
greedy = argmax of the logits at every step. TheT → 0limit. Simple, deterministic, often boring. - Top-k sampling —
top-k = keep the k highest-probability tokens, sample from those. A truncation, not a reshape. - Top-p / nucleus sampling —
top-p = keep tokens whose probabilities sum to ≥ p, sample from those. Adapts how many tokens are in play based on how peaked the distribution already is. Often combined with temperature. - Logit bias —
logit bias = additive nudge to a specific token's logit before sampling— the “I want to forbid this token” or “I want to encourage that one” knob. - LLM —
LLM = transformer + next-token objective at scale— the thing producing the logits in the first place. - Hallucination —
hallucination = confidently wrong output— interacts with temperature but is not caused by it.
Going deeper
- The Curious Case of Neural Text Degeneration (Holtzman et al., 2019) — the primary source for “why is the most likely continuation the wrong thing to always pick?”, and where nucleus sampling was introduced.
- Hugging Face’s How to generate text — the explainer for “how do greedy, beam, top-k, top-p, and temperature actually differ?”, with runnable code you can poke at.
- Your provider’s own API reference, read side by side with another’s:
answers “what does this number actually do here?”, since the scales,
the defaults, and the order of composition with
top_p/top_kdiffer between vendors and are the usual reason a prompt behaves differently after you switch models. - Rabbit hole: the
Boltzmann distribution
in statistical physics — for “why is it called temperature at all?”: the
exp(−E/T)form is the same one, with highTspreading particles across states and lowTsettling them into the ground state.
The math here (logit / T inside softmax), the qualitative behavior across the
Trange, and the fact that “temperature 0” on hosted APIs is near-deterministic but not bit-exact are all well-established. What is not public is the order in which any specific provider composes temperature with top-p / top-k, and whatever internal logit renormalization it does before exposing the knob — providers rarely document either, and both change across model versions. Check the current API reference rather than a blog post.