Why isn't temperature 0 actually deterministic?
You set temperature to 0, send the same prompt twice, get two different answers. The math says argmax is a function. The hardware disagrees.
On this page
The picture version
Five pictures for a reader whose regression test just flaked. The prose below fills in the seams the pictures skip.
1 · The problem
Same prompt, temperature 0, two different answers.
2 · Layer one
Adding the same numbers in a different order gives a different answer.
3 · Layer two
You don’t choose the order. A kernel library does, at runtime.
4 · Layer three
Your batch isn’t yours.
5 · Keep this card
The whole thing on one index card.
Why it exists
Every engineer who has tried to write a regression test against a hosted LLM hits this wall.
You set temperature: 0. You re-send the exact same prompt. You expect the
exact same output, because the docs say temperature 0 means “always pick the
most likely token” and that’s a function — same input, same output, end of
story. Then the second response comes back slightly different. Sometimes a
word, sometimes a whole reordered paragraph. Run it ten times and you’ll see
two or three variants.
This is not a bug in your code. It’s not the model being “creative.” The sampling step really is argmax. The reason the output drifts has nothing to do with the model’s distribution and everything to do with what’s happening to the numbers underneath: floating-point arithmetic on a GPU, running in a kernel whose behavior depends on what other requests happen to be in the same batch as yours.
It’s worth understanding because the determinism story breaks in a place people don’t look. If you assume “temperature 0 = reproducible,” you’ll write tests that flake, caches that miss, and evals that drift between runs of the same model on the same hardware. The model isn’t lying to you. The abstraction is.
Why it matters now
Determinism is load-bearing for a lot of what people are trying to build on top of LLMs in 2026:
- Evals and benchmarks. “We re-ran the same eval and the score moved by half a point” is constant noise in the model-comparison world. Some of that is genuine model variance. A surprising amount is just nondeterministic inference on the same model.
- Caching by output. You can store one observed response, but you can’t assume a later call regenerates the same bytes, so “this prompt → this response” isn’t a safe equivalence to build on. (Caching by prompt still works, which is what providers actually ship.)
- Agent debugging. When a coding agent does the wrong thing, the first thing you want is to replay the run. If the same prompt at T=0 gives a different tool call, you can’t isolate whether the bug was in the model, the harness, or the world it was acting on.
- Scientific reproducibility. Papers that report a benchmark number without specifying batch size, hardware, kernel library, and concurrency conditions are reporting a number with hidden error bars.
Recent engineering work (see Going deeper) has shown that the user-visible piece of this problem is mostly tractable if you make inference kernels batch-invariant. But the default for hosted APIs and most open-source serving stacks is still “nearly deterministic, not bit-exact.” The right mental model is the latter.
The short answer
nondeterminism at T=0 = floating-point non-associativity + batch-dependent GPU kernels
Picture to keep: a photo finish. Two tokens cross the line a thousandth of a nose apart, and the camera you’re judging with gets re-mounted at a slightly different angle every race — not because anyone touched your race, but because the other runners on the track changed which camera rig the stadium set up. Same runners, same speeds, different call.
Argmax over the logits is deterministic. The logits themselves aren’t bit-identical across runs, because the matrix multiplications that produce them sum thousands of floating-point numbers in a different order each time — and on top of that, the kernels chosen depend on the shape of the batch your request happens to land in. Different order of additions, different rounding, occasionally different argmax.
How it works
Follow what happens as you keep trying to fix it, because each fix
reveals the next layer down. You started by setting temperature: 0 —
that was the obvious move, and it closed the sampling question and
nothing else. Below the sampler there are three more layers, and each
one is enough on its own to break bit-exact reproducibility.
1. Floating-point addition isn’t associative
In real arithmetic, (a + b) + c = a + (b + c). In
floating-point
arithmetic, those two expressions can give different answers, because each
intermediate sum is rounded to fit in 16 or 32 bits. Sum a thousand
near-zero numbers in one order, get one result. Sum them in a different
order, get a result that differs in the last few bits.
A single attention layer’s logits come from summing thousands of products. The final logits — the ones argmax runs on — are the result of many such reductions stacked through dozens of layers. Tiny rounding differences anywhere in that chain can flip a near-tie at the top of the final distribution.
Most of the time, the top token is far enough ahead that this doesn’t
matter. Occasionally, two tokens are within a hair’s breadth — say, the
at logit 8.4012 and a at 8.4007 — and a different summation order tips
which one wins. That single token then changes the rest of the generation,
because the model conditions on its own outputs.
2. GPU kernels can reduce in input-shape-dependent order
So the next fix suggests itself: fix the summation order and the rounding stops varying. That’s harder than it sounds, because you don’t choose the summation order — a kernel library does, at runtime.
You could in principle write a matmul that always sums in the same order for a given input shape. Real high-performance GPU kernels do something weaker: they’re often deterministic given a fixed input shape, but the reduction strategy they pick can vary with shape. Sources of variability:
- Auto-tuned kernel selection. Libraries like cuBLAS and cuDNN pick from multiple kernel implementations at runtime based on tensor shapes; two different shapes can pick two different algorithms with different tiling and reduction trees.
- Parallel reduction trees. A sum of N numbers is split across many threads; the partial-sum tree’s shape depends on tile and split choices, which depend on shape.
- Atomic adds in some routines. Certain GPU operations accumulate into shared memory with atomics, which complete in scheduling order. NVIDIA documents this for specific routines. Worth noting: for typical LLM forward passes, atomics aren’t usually the operative cause of run-to-run variation — the Thinking Machines piece linked below is pretty firm on this point.
The major frameworks expose a “deterministic mode” flag that forces order-stable kernels — PyTorch documents its controls in the reproducibility notes. Turning it on costs throughput. Hosted inference providers, optimizing for tokens/sec/dollar, don’t turn it on by default.
3. Your batch isn’t your batch
Fine — pin the kernels, then. Deterministic mode on, fixed input shape, same hardware, same weights. Everything about your request is now nailed down. And on a hosted API it still drifts, because the input shape was never yours to fix.
This is the one most engineers don’t see coming, and the recent Thinking Machines piece (below) argues it’s the dominant cause of T=0 drift on hosted APIs.
Inference servers don’t run one request at a time. They batch many concurrent requests together so the GPU isn’t sitting idle. When your request arrives, it gets stitched into a batch with whoever else is on the server right now: a 200-token prompt next to a 10,000-token prompt, padded or packed together, processed in a single forward pass.
The mechanism isn’t that other users’ values leak into yours — they don’t. It’s that the shape of the combined tensor changes which kernel implementation gets picked, which tiling it uses, and therefore the order in which your own values get summed. The arithmetic on your numbers changes because they’re sitting in a differently-shaped tensor.
For mixture-of-experts models, there’s an additional path: in capacity-limited MoE implementations, expert capacity is allocated per batch, and two requests competing for the same expert can cause one to overflow and get routed differently than it would alone. Whether this actually happens depends on the specific MoE serving policy.
So you can send the same prompt twice in a row, and the only thing that changed between call 1 and call 2 is which other users were on the server. That’s enough to change kernel selection and occasionally flip the argmax.
This is why local inference (one model, one request, fixed batch shape, fixed hardware, deterministic kernels enabled) can be made bit-reproducible, and hosted APIs typically aren’t — not without engineering work that gives up some of the batching flexibility that makes them affordable.
Where it gets subtle
- It’s usually small. Typically a prompt runs identically for a while, then diverges once a near-tie hits. The output is qualitatively the same; the bytes aren’t.
- Greedy decoding hides it less than you’d think. People assume that
setting
temperature: 0and calling it “deterministic mode” closes the question. It only closes the sampling question. Everything upstream of the sampler is still nondeterministic. - Different precisions amplify it differently. Models served in bfloat16 or fp8 have less headroom against rounding noise than fp32. Quantized serving can make near-ties flip more often.
- “Seeded” APIs don’t solve it. Some providers offer a
seedparameter, and the docs themselves usually say it’s best-effort and not guaranteed reproducible. The seed pins the sampler’s randomness — useful at T > 0 — but the upstream kernel and batch-shape variation is still there. Same seed, same prompt, T=0 on a busy server: still drifts. - It can be defeated, with effort. It is possible to build a fully deterministic inference stack — pinned kernels, fixed batch shapes, invariant reductions across batch sizes. The cost is throughput and engineering work. Recent research has shown that most of the nondeterminism people attribute to GPUs is really about batch-invariance of kernels, and that fixing that makes inference reproducible at a real but bounded throughput cost.
You started with nondeterminism at T=0 = floating-point non-associativity + batch-dependent GPU kernels. What did chasing the
fixes add? — + the batch isn't yours. Non-associativity alone would
be a solved problem; you’d pin the reduction order and go home. The
part that makes hosted inference irreducibly non-reproducible is that
the tensor shape your arithmetic runs inside is decided by who else
was on the server, which is not a parameter you can pass. Argmax is
still a function; you just never send it the same input twice.
Check yourself
Before you go — you move a flaky eval off the hosted API onto your own
box: one GPU, one request at a time, temperature: 0, framework
deterministic mode on, same model weights. Should you now expect
byte-identical output across runs? And what would you still not
expect to match?
Answer
Yes, that’s the configuration where bit-exact reproducibility becomes achievable, and it’s achievable precisely because you removed the third mechanism — nobody else’s request is reshaping your tensors, so the batch shape is constant, so kernel selection is constant, so the reduction order is constant, so the logits are bit-identical and argmax is a function again. What you should not expect to survive is a change to any of those constants: a different GPU model, a different CUDA or kernel-library version, a different batch size or max-tokens setting, a different quantization, or turning deterministic mode back off. Reproducibility here is a property of the whole stack, not of the model — which is why a benchmark number reported without that stack described has hidden error bars.
And one more — a colleague argues the drift can’t matter much, since the model’s actual preferences aren’t changing, only the last few bits of the logits. Where does that reasoning fail?
Answer
It’s right about the magnitude and wrong about the consequence, because
generation is autoregressive. A last-bit difference is harmless on the
99%+ of steps where the top token wins comfortably. But when two tokens
are within a hair — the at 8.4012 versus a at 8.4007 — the tie
breaks the other way, and from that point on the model is conditioning
on a different prefix. The divergence isn’t proportional to the
numerical error; a single flipped token can send the rest of the
response down an entirely different path. That’s why you typically see
identical output for the first N tokens and then a fork, rather than
uniformly slightly-different text.
Famous related terms
- Temperature —
temperature = a knob that flattens or sharpens the logit distribution before sampling. The thing people think controls determinism. - Greedy decoding —
greedy = argmax of the logits at every step. Deterministic on identical logits; the logits are the problem. - KV cache —
KV cache = stored keys and values for past tokens, reused on each step. Cache layout and paging strategy can subtly change reduction order too. - Batch invariance —
batch invariance = a kernel returns bit-identical results regardless of how its inputs are batched. The property that, if enforced everywhere, would make hosted inference reproducible. - Floating-point non-associativity —
(a + b) + c ≠ a + (b + c)in finite precision. The root cause sitting under everything else. - Seeded sampling —
seed = pin the sampler's RNG. Fixes nondeterminism at T > 0; does nothing at T = 0. - LLM — the thing whose logits we’re arguing about.
Going deeper
- Defeating Nondeterminism in LLM Inference — Horace He and collaborators at Thinking Machines Lab, September 2025. An engineering note (not a peer-reviewed paper) arguing that the user-visible source of T=0 drift on hosted LLMs is specifically the batch-invariance of inference kernels — i.e. that kernels return different results for the same logical input depending on how they’re batched — rather than generic GPU scheduling noise. Sharper than this post on the actual mechanism; worth reading.
- PyTorch’s
reproducibility docs
— concrete on which operations are nondeterministic on GPU and what
torch.use_deterministic_algorithms(True)actually changes. - What Every Computer Scientist Should Know About Floating-Point Arithmetic (David Goldberg, ACM Computing Surveys 23(1), March 1991) — the canonical reference for why the arithmetic doesn’t behave the way the math does. Thirty-five years old and still the right starting point.
Well-established: the three mechanisms themselves — floating-point non-associativity, shape-dependent kernel selection, and batch composition affecting that shape. Contested: their relative weight. The Thinking Machines piece in Going deeper argues batch-invariance of kernels is the dominant user-visible cause on hosted LLMs, and vLLM’s own reproducibility docs point the same way, but no published measurement compares the three across providers. Not public: the provider internals that would settle it — kernel libraries, batching policy, serving precision, MoE routing. So read the three as “all real, and they stack,” not as a settled ranking for any one stack.