Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why floating-point addition isn't associative

Schoolroom math says (a + b) + c equals a + (b + c). On a real computer it doesn't, and that one fact ripples out into nondeterministic GPU reductions, irreproducible training runs, and LLM outputs that aren't bit-stable across hardware.

Computer Science intermediate May 2, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has never thought about how a computer stores a number. The prose below fills in the seams the pictures skip.

1 · The problem

Same script, same seed, same machine. Slightly different answer.

the version everyone has seen 0.1 + 0.2 0.30000000000000004 the version that costs you a day run the identical script twice loss differs at step 10,000 Nothing changed. Not the code, not the seed, not the data. One unintuitive fact explains both: the order you add numbers in changes the answer
Everyone meets the small version and files it under “rounding.” The consequence worth knowing isn’t the odd digit at the end — it’s that addition stops obeying a rule you have assumed since school.

2 · The counterexample

Three numbers. Two groupings. Two different answers.

a = 1e20   b = −1e20   c = 1 (a + b) + c the giants cancel exactly → 0 then 0 + 1 = 1 a + (b + c) the 1 is added to a giant first and disappears into it = 0 Same three numbers. Same operator. Only the brackets moved. and both answers are correct by the rules the hardware actually follows
This is the whole phenomenon in three numbers, and you can reproduce it in any language in one line. The 1 survives or vanishes depending purely on what it was added next to — which the bracketing decides.

3 · Why the 1 can vanish

The gaps between storable numbers grow as the numbers do.

near zero: neighbours are a whisker apart out here: neighbours are trillions apart every value a computer can actually store, marked on the line Adding 1 out there asks for a value between two neighbours. There isn’t one — so the answer rounds straight back to where it started.
A fixed number of bits can only mark so many points on the number line, and they are spaced proportionally rather than evenly. Out at enormous magnitudes the spacing dwarfs 1, so adding 1 changes nothing at all.

4 · The rule that follows

Every step rounds. So the order decides what gets rounded away.

two numbers the exact answer nearest storable round That happens after every single operation, not once at the end. Change the order and you change which intermediate values exist — and therefore which ones get rounded, and by how much. this is not a bug anyone forgot to fix: it follows from defining each operation exactly
Each operation is specified as “work out the exact answer, then round it to the nearest storable value.” Rounding loses information, and which information you lose depends on the order — so the loss of associativity is forced by the definition rather than overlooked.

5 · Why this shows up on a graphics card

A million numbers aren’t added left to right. They’re added as a tree.

how you’d write it one after another, always the same order how the hardware does it in pairs, in parallel — and the shape of the tree can differ between runs of the same code Different tree, different intermediate values, different final answer.
Parallel hardware adds in a tree of partial sums whose shape depends on the sizes involved and the routine the runtime happened to pick. Two runs of the same code can associate differently, which is why an identical script can produce a loss curve that is close but not identical.

6 · Keep this card

The whole thing on one index card.

why order matters = a fixed number of digits + rounding after every operation ∴ re-order the sums, re-order what gets lost What you gain is that every single operation is exactly specified. What you give up is the algebra you learned at school.
Picture to keep: a scale whose resolution gets coarser the more weight is already on it — a paperclip registers on an empty scale and not on a loaded one, so whether a small quantity survives depends on what it was added next to. Where it breaks: a real scale rounds the reading, while a computer rounds the stored value itself, so the paperclip isn’t sitting there unmeasured. It is gone.

Why it exists

You have almost certainly seen the small version of this. Open any Python prompt, or a JavaScript console, or a spreadsheet cell, and add 0.1 + 0.2. You get 0.30000000000000004. Everyone’s first reaction is “the computer got arithmetic wrong,” and everyone’s first explanation is “rounding.” Both are close enough — but the interesting consequence of that rounding isn’t a weird digit at the end. It’s this: the order you add numbers in changes the answer.

Which brings us to the version that costs people real time. You re-run the same training script with the same seed, the same data, the same model, the same hardware. The loss curve at step 10,000 is almost the same as last time — but not quite. Off by 0.0003. Run it again: off by something different. Nothing in your code changed. Nothing in your config changed. The bits going into the GPU are identical. The bits coming out aren’t. This is not a bug in CUDA or in your framework. It’s the consequence of a single, deeply unintuitive fact about computer arithmetic: floating-point addition isn’t associative. (a + b) + c and a + (b + c) can give different answers, and on modern hardware they routinely do.

The reason is finite precision. A real number has, in principle, infinitely many digits. A floating-point number has a fixed number of bits — 32, 16, 8 — split between an exponent (how big) and a mantissa (how precisely). After every single arithmetic operation, the result is rounded back to fit. Round, round, round. Once you accept that, the non-associativity is forced: which intermediate values you produce — and therefore which ones get rounded — depends on the order you produced them in. Different order, different intermediates, different rounding errors, different final answer.

One three-number example carries this whole post: a = 1e20, b = -1e20, c = 1. Before reading on, commit to a guess — does (a + b) + c equal a + (b + c), and if not, which one loses the 1?

This is not a flaw the standards committee forgot to fix. IEEE 754 — the spec almost every CPU and GPU implements — defines each operation as “compute the exact result, then round it to the nearest representable value.” Non-associativity follows from that definition rather than being legislated separately, and the committee accepted the consequence rather than missing it. What you get in exchange is that every individual operation is exactly specified and portable; what you give up is the algebraic laws of the real numbers.

Why it matters now

This sounds like a textbook curiosity. It isn’t — it’s the load-bearing explanation for several things engineers hit constantly.

The short answer

float non-associativity = finite mantissa + rounding after every op

Picture to keep: a scale whose resolution gets coarser the more weight is already on it. Step on holding a paperclip and the reading doesn’t budge — the paperclip is finer than the tick marks at your weight. Put the paperclip on an empty scale and it registers fine. Whether a small quantity survives depends on what it was added next to, and the order of operations decides that. (Where the analogy breaks: a real scale rounds readings; a float rounds the stored value itself, so the loss is permanent — the paperclip isn’t sitting there un-measured, it’s gone.)

Real numbers carry as many digits as they need. Floating-point numbers don’t — every result gets rounded to a fixed mantissa width. Rounding is where the information loss happens, and information loss isn’t order-invariant. Re-arrange the additions, you re-arrange which intermediate values get rounded, and that changes the final answer. The same equation, evaluated two valid ways, gives two slightly different floats.

How it works

Here is the running example, run for real. In FP32, with a = 1e20, b = -1e20, c = 1:

import numpy as np
a, b, c = np.float32(1e20), np.float32(-1e20), np.float32(1.0)
(a + b) + c   # → 1.0
a + (b + c)   # → 0.0

Same three numbers. Same operator. Different parenthesization. Different answer. (You can paste those lines into a Python shell and reproduce the split yourself.)

Walk through (a + b) + c. First 1e20 + (-1e20) = 0 exactly — two equal- magnitude opposites cancel, no rounding needed. Then 0 + 1 = 1. The 1 survived because it never got near a giant number.

Walk through a + (b + c). First -1e20 + 1. The number -1e20 lives at a magnitude where the gap between consecutive FP32 values is enormous — much bigger than 1. (That gap has a name: ULP.) At 1e20, one ULP in FP32 is roughly 2^43 ≈ 9e12. Adding 1 to -1e20 produces a true result that lies between two representable FP32 values, and rounding picks the nearer one — which is -1e20 itself. The 1 was annihilated by the rounding step. Then 1e20 + (-1e20) = 0. The 1 is gone, never to return.

Re-association is what decided whether the small number survived. Pair it with its near-equal partner first and it lives. Pair it with a giant first and it’s silently rounded away.

The general shape: every floating-point operation is round(true_result). Re-association doesn’t change true_result, but it changes which intermediate gets rounded. If a small number meets a big number first, the small one is silently swallowed. If two big numbers meet first and cancel, the small one survives. The associative law is a property of the exact arithmetic underneath, not of the rounded arithmetic the hardware actually performs.

A few consequences fall out of this:

The honest seam: I’m describing the mechanism — finite mantissa plus per-op rounding plus reordering. The empirical question of how much this matters for a given workload (training a 70B model? doing a 3×3 matmul?) depends on numerical conditioning, precision, and reduction strategy in ways that don’t compress to a single rule. For a deeper treatment, the canonical reference is Goldberg, What Every Computer Scientist Should Know About Floating-Point Arithmetic (1991).

You started with float non-associativity = finite mantissa + rounding after every op. What did 1e20 + (-1e20) + 1 add to that line? — + re-grouping changes which value gets rounded, not what the true answer is. That’s why the fix is never “use more precision” (a wider mantissa just moves the threshold); it’s “control the order,” which is exactly what deterministic-mode flags buy you and exactly what they cost throughput for.

Check yourself

Before you go — you sum a million positive floats. Your colleague sorts them smallest-to-largest first and gets a different (and closer to exact) total. Why does sorting help, and would sorting largest-to-smallest help too?

Answer

Summing ascending keeps the running total small for as long as possible, so each new addend is comparable in magnitude to the accumulator and fewer bits get rounded away — the paperclip meets the paperclips, not the person. Sorting descending does the opposite: the accumulator becomes huge immediately and every subsequent small value risks being swallowed. Note the honest limit — ascending order is a heuristic, not a guarantee, and Kahan summation beats both without needing a sort.

And: your inference server returns bit-identical logits for a prompt when you send it alone, but slightly different ones when the same prompt shares a batch with other requests. Is this a bug in the model weights?

Answer

No — the weights are identical. The most commonly cited mechanism is that batch size changes the shapes going into the matmuls, which changes which kernel and tile decomposition the library picks, which changes how the partial sums get grouped. (It isn’t the only possible mechanism — atomics and backend selection can do it too — but the shape is always the same: different grouping, different rounding, slightly different logits.) Every one of those answers is a valid IEEE 754 computation; none is “the correct” one. If a top-2 token pair is near-tied, that drift is enough to flip the argmax.

Going deeper