Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why isn't temperature 0 actually deterministic?

You set temperature to 0, send the same prompt twice, get two different answers. The math says argmax is a function. The hardware disagrees.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Five pictures for a reader whose regression test just flaked. The prose below fills in the seams the pictures skip.

1 · The problem

Same prompt, temperature 0, two different answers.

temperature: 0 the exact same prompt, twice “The function returns the parsed config object.” “This function returns a parsed config object.” run it ten times and you will see two or three variants The sampling step really is argmax. Argmax really is a function. so this is not the model being creative, and not a bug in your code The drift is in the numbers argmax is handed.
If you assume “temperature 0 = reproducible” you will write tests that flake, caches that miss, and evals that drift between runs of the same model on the same hardware. The model isn’t lying to you. The abstraction is.

2 · Layer one

Adding the same numbers in a different order gives a different answer.

in real arithmetic these are equal. in floating point they are not. (a + b) + c a + (b + c) every intermediate sum is rounded to fit in 16 or 32 bits so at the top of the final distribution, a near-tie can tip either way: the   8.4012 a     8.4007 a different summation order moves the last few bits — and occasionally that is enough to swap which one wins One flipped token changes everything after it, because the model reads its own output.
The final logits are the result of many stacked reductions, each summing thousands of products across dozens of layers. Most of the time the top token is far enough ahead that none of this matters — and then occasionally it isn’t.

3 · Layer two

You don’t choose the order. A kernel library does, at runtime.

your tensor, some shape the kernel library picks an implementation based on that shape one tiling, one reduction tree another tiling, another tree a sum of N numbers is split across many threads; the shape of the partial-sum tree follows the tiling Frameworks ship a “deterministic mode” that forces order-stable kernels. it costs throughput, so hosted providers optimising for tokens per second per dollar leave it off Fine. Turn it on, pin the shape, and you are still not done.
High-performance kernels are often deterministic given a fixed input shape — the variability comes from the shape changing which algorithm gets selected. Which sets up the layer nobody sees coming.

4 · Layer three

Your batch isn’t yours.

call 1 — who else happened to be on the server you call 2 — same prompt, same weights, same hardware you a different combined shape → a different kernel → a different summation order nobody else’s values leak into yours. the arithmetic on your numbers changes because they sit in a differently-shaped tensor. The one input you cannot pin is the one you never sent. which is why local single-request inference can be made bit-reproducible and a busy hosted API generally cannot
Servers batch concurrent requests so the GPU isn’t idle — a 200-token prompt packed next to a 10,000-token one, processed in a single forward pass. The only thing that changed between your two calls was who else was on the server.

5 · Keep this card

The whole thing on one index card.

nondeterminism at T=0 = floating-point non-associativity + batch-dependent GPU kernels ∴ argmax is still a function — you just never send it    the same input twice
Picture to keep: a photo finish. Two tokens cross the line a thousandth of a nose apart, and the camera you’re judging with gets re-mounted at a slightly different angle every race — not because anyone touched your race, but because the other runners on the track changed which rig the stadium set up. Same runners, same speeds, different call.

Why it exists

Every engineer who has tried to write a regression test against a hosted LLM hits this wall.

You set temperature: 0. You re-send the exact same prompt. You expect the exact same output, because the docs say temperature 0 means “always pick the most likely token” and that’s a function — same input, same output, end of story. Then the second response comes back slightly different. Sometimes a word, sometimes a whole reordered paragraph. Run it ten times and you’ll see two or three variants.

This is not a bug in your code. It’s not the model being “creative.” The sampling step really is argmax. The reason the output drifts has nothing to do with the model’s distribution and everything to do with what’s happening to the numbers underneath: floating-point arithmetic on a GPU, running in a kernel whose behavior depends on what other requests happen to be in the same batch as yours.

It’s worth understanding because the determinism story breaks in a place people don’t look. If you assume “temperature 0 = reproducible,” you’ll write tests that flake, caches that miss, and evals that drift between runs of the same model on the same hardware. The model isn’t lying to you. The abstraction is.

Why it matters now

Determinism is load-bearing for a lot of what people are trying to build on top of LLMs in 2026:

Recent engineering work (see Going deeper) has shown that the user-visible piece of this problem is mostly tractable if you make inference kernels batch-invariant. But the default for hosted APIs and most open-source serving stacks is still “nearly deterministic, not bit-exact.” The right mental model is the latter.

The short answer

nondeterminism at T=0 = floating-point non-associativity + batch-dependent GPU kernels

Picture to keep: a photo finish. Two tokens cross the line a thousandth of a nose apart, and the camera you’re judging with gets re-mounted at a slightly different angle every race — not because anyone touched your race, but because the other runners on the track changed which camera rig the stadium set up. Same runners, same speeds, different call.

Argmax over the logits is deterministic. The logits themselves aren’t bit-identical across runs, because the matrix multiplications that produce them sum thousands of floating-point numbers in a different order each time — and on top of that, the kernels chosen depend on the shape of the batch your request happens to land in. Different order of additions, different rounding, occasionally different argmax.

How it works

Follow what happens as you keep trying to fix it, because each fix reveals the next layer down. You started by setting temperature: 0 — that was the obvious move, and it closed the sampling question and nothing else. Below the sampler there are three more layers, and each one is enough on its own to break bit-exact reproducibility.

1. Floating-point addition isn’t associative

In real arithmetic, (a + b) + c = a + (b + c). In floating-point arithmetic, those two expressions can give different answers, because each intermediate sum is rounded to fit in 16 or 32 bits. Sum a thousand near-zero numbers in one order, get one result. Sum them in a different order, get a result that differs in the last few bits.

A single attention layer’s logits come from summing thousands of products. The final logits — the ones argmax runs on — are the result of many such reductions stacked through dozens of layers. Tiny rounding differences anywhere in that chain can flip a near-tie at the top of the final distribution.

Most of the time, the top token is far enough ahead that this doesn’t matter. Occasionally, two tokens are within a hair’s breadth — say, the at logit 8.4012 and a at 8.4007 — and a different summation order tips which one wins. That single token then changes the rest of the generation, because the model conditions on its own outputs.

2. GPU kernels can reduce in input-shape-dependent order

So the next fix suggests itself: fix the summation order and the rounding stops varying. That’s harder than it sounds, because you don’t choose the summation order — a kernel library does, at runtime.

You could in principle write a matmul that always sums in the same order for a given input shape. Real high-performance GPU kernels do something weaker: they’re often deterministic given a fixed input shape, but the reduction strategy they pick can vary with shape. Sources of variability:

The major frameworks expose a “deterministic mode” flag that forces order-stable kernels — PyTorch documents its controls in the reproducibility notes. Turning it on costs throughput. Hosted inference providers, optimizing for tokens/sec/dollar, don’t turn it on by default.

3. Your batch isn’t your batch

Fine — pin the kernels, then. Deterministic mode on, fixed input shape, same hardware, same weights. Everything about your request is now nailed down. And on a hosted API it still drifts, because the input shape was never yours to fix.

This is the one most engineers don’t see coming, and the recent Thinking Machines piece (below) argues it’s the dominant cause of T=0 drift on hosted APIs.

Inference servers don’t run one request at a time. They batch many concurrent requests together so the GPU isn’t sitting idle. When your request arrives, it gets stitched into a batch with whoever else is on the server right now: a 200-token prompt next to a 10,000-token prompt, padded or packed together, processed in a single forward pass.

The mechanism isn’t that other users’ values leak into yours — they don’t. It’s that the shape of the combined tensor changes which kernel implementation gets picked, which tiling it uses, and therefore the order in which your own values get summed. The arithmetic on your numbers changes because they’re sitting in a differently-shaped tensor.

For mixture-of-experts models, there’s an additional path: in capacity-limited MoE implementations, expert capacity is allocated per batch, and two requests competing for the same expert can cause one to overflow and get routed differently than it would alone. Whether this actually happens depends on the specific MoE serving policy.

So you can send the same prompt twice in a row, and the only thing that changed between call 1 and call 2 is which other users were on the server. That’s enough to change kernel selection and occasionally flip the argmax.

This is why local inference (one model, one request, fixed batch shape, fixed hardware, deterministic kernels enabled) can be made bit-reproducible, and hosted APIs typically aren’t — not without engineering work that gives up some of the batching flexibility that makes them affordable.

Where it gets subtle

You started with nondeterminism at T=0 = floating-point non-associativity + batch-dependent GPU kernels. What did chasing the fixes add? — + the batch isn't yours. Non-associativity alone would be a solved problem; you’d pin the reduction order and go home. The part that makes hosted inference irreducibly non-reproducible is that the tensor shape your arithmetic runs inside is decided by who else was on the server, which is not a parameter you can pass. Argmax is still a function; you just never send it the same input twice.

Check yourself

Before you go — you move a flaky eval off the hosted API onto your own box: one GPU, one request at a time, temperature: 0, framework deterministic mode on, same model weights. Should you now expect byte-identical output across runs? And what would you still not expect to match?

Answer

Yes, that’s the configuration where bit-exact reproducibility becomes achievable, and it’s achievable precisely because you removed the third mechanism — nobody else’s request is reshaping your tensors, so the batch shape is constant, so kernel selection is constant, so the reduction order is constant, so the logits are bit-identical and argmax is a function again. What you should not expect to survive is a change to any of those constants: a different GPU model, a different CUDA or kernel-library version, a different batch size or max-tokens setting, a different quantization, or turning deterministic mode back off. Reproducibility here is a property of the whole stack, not of the model — which is why a benchmark number reported without that stack described has hidden error bars.

And one more — a colleague argues the drift can’t matter much, since the model’s actual preferences aren’t changing, only the last few bits of the logits. Where does that reasoning fail?

Answer

It’s right about the magnitude and wrong about the consequence, because generation is autoregressive. A last-bit difference is harmless on the 99%+ of steps where the top token wins comfortably. But when two tokens are within a hair — the at 8.4012 versus a at 8.4007 — the tie breaks the other way, and from that point on the model is conditioning on a different prefix. The divergence isn’t proportional to the numerical error; a single flipped token can send the rest of the response down an entirely different path. That’s why you typically see identical output for the first N tokens and then a fork, rather than uniformly slightly-different text.

Going deeper

Well-established: the three mechanisms themselves — floating-point non-associativity, shape-dependent kernel selection, and batch composition affecting that shape. Contested: their relative weight. The Thinking Machines piece in Going deeper argues batch-invariance of kernels is the dominant user-visible cause on hosted LLMs, and vLLM’s own reproducibility docs point the same way, but no published measurement compares the three across providers. Not public: the provider internals that would settle it — kernel libraries, batching policy, serving precision, MoE routing. So read the three as “all real, and they stack,” not as a settled ranking for any one stack.