Why bf16 won the training format wars
Half-precision floats came in two flavors: fp16, which had been around for years, and bf16, which kept fp32's exponent and threw away mantissa bits. The less-precise format won. Here's why that's not a typo.
On this page
The picture version
Six pictures for a reader who has never thought about how many digits a number gets. The prose below fills in the seams the pictures skip.
1 · The problem
Same sixteen bits, split two ways. One of them loses your number entirely.
2 · What the bits are for
Some bits say how big. The rest say how precisely.
3 · What reach buys
Two rulers with the same number of ticks. One is short.
4 · What reach costs
Wider ruler, coarser ticks. About three decimal digits, and that’s your lot.
5 · Why the narrow one came first
It was designed for pictures, where nothing is ever that small.
6 · Keep this card
The whole thing on one index card.
Why it exists
Punch 0.0000001 × 0.0000001 into your phone’s calculator. It doesn’t
print a long crawl of zeros and it doesn’t give up — it shows something
like 1e-14. The display has a fixed number of character slots, and
scientific notation spends a few of them on the exponent so the rest can
carry digits. You just traded precision for range without thinking about
it, and you got a number you could still read.
That’s the only decision a 16-bit float format ever makes: given a fixed bit budget, how much goes to how big or small a number you can write down and how much to how precisely you can write it. fp16 spent more on precision. bf16 spent more on range. For training neural networks, range turned out to matter far more — which is why the less precise format won.
Keep one number in your pocket for the rest of this post: a single
weight gradient that comes out to 1e-9. Everything below is the story
of whether that number survives the trip from the backward pass into the
optimizer.
Here’s the puzzle. You’re training a neural network. You’d like to use 16-bit floats instead of 32-bit floats — half the memory, half the bandwidth, often more than double the throughput on modern matrix engines. There are two candidates sitting on the shelf:
- fp16, standardized in IEEE 754, has been in graphics hardware for over a decade. 5 exponent bits, 10 mantissa bits.
- bf16, invented at Google for the TPU and later adopted by everyone else. 8 exponent bits, 7 mantissa bits.
Read those numbers carefully. bf16 has fewer mantissa bits than fp16 — 7 vs. 10. It’s the less precise of the two. And yet bf16 is now the default mixed-precision format on every major training stack I’ve seen public recipes for, and fp16 is treated as the format you have to work around.
That should read as backwards, because the entire historical case for floating point was “trade range for precision smartly, but get plenty of both.” bf16 deliberately throws precision away to keep range. Why was that the right trade?
The short version: training a neural network is a lot less sensitive to how precisely you can represent a number than to how small or large a number you can represent at all. fp16’s narrower exponent squeezes from both ends — small gradients underflow to zero, and scaled-up values can overflow to infinity — so engineers spent years inventing workarounds (loss scaling, mixed-precision recipes) to keep training inside the window. bf16 just gives the gradients somewhere to live, accepts that the last couple of bits of the mantissa are noise anyway, and gets out of the way.
Why it matters now
If you train models, this is one of the few hardware decisions that’s still
visible in your code. PyTorch’s torch.bfloat16 and torch.float16 are not
interchangeable, and picking the wrong one is the difference between a stable
run and one that quietly diverges to NaN partway through.
It also explains why specific accelerators became popular when they did. Google developed bfloat16 for Cloud TPU v2 and v3; Cloud TPU went publicly available in beta in February 2018. NVIDIA’s tensor cores gained native bf16 with the A100 (Ampere, 2020). After that point, “use mixed precision” stopped meaning “fp16 with a bag of tricks” and started meaning “bf16, which Just Works.” Frameworks shifted defaults, and by the H100 generation (2022) bf16 was the common choice in publicly described large-model training recipes running on bf16-capable hardware.
It matters even outside training. Inference is moving toward fp8 and below (fp8 in H100, fp4 announced for Blackwell), and the design decisions there inherit bf16’s lesson: when you have to throw bits away, throw away precision before you throw away range.
The short answer
bf16 = sign + 8-bit exponent (same as fp32) + 7-bit mantissa
Picture to keep: two rulers with the same number of tick marks. fp16’s
ruler is short but finely etched; bf16’s stretches from the atomic to the
astronomical with coarse ticks. Our 1e-9 gradient falls off the end of
fp16’s ruler entirely, and lands comfortably in the middle of bf16’s — a
coarse tick is still a tick. (Where the ruler breaks: a float’s ticks aren’t
evenly spaced like a real ruler’s. They’re spaced proportionally, so the
gaps grow with the magnitude — which is why big numbers swallow small ones.)
bf16 is essentially fp32 with the bottom 16 bits of the mantissa lopped off (plus a rounding rule, and on some hardware a flush-to-zero treatment of denormals). It has the same exponent range as fp32 — the same ability to represent very small gradients and very large activations — but only enough mantissa precision for roughly 2–3 significant decimal digits. That’s plenty for gradient descent, which spends its life adding noisy small updates to noisy current weights. It is not plenty for, say, a stiff differential equation solver — but neural network training isn’t that.
How it works
The oddity to keep in view is that the less precise format won, and the
reason is that the exponent and the mantissa fail in completely different
ways. A floating-point number is sign × mantissa × 2^exponent. You get a
fixed bit budget; you split it between mantissa and exponent.
fp32: 1 sign | 8 exponent | 23 mantissa (32 bits total)
fp16: 1 sign | 5 exponent | 10 mantissa (16 bits)
bf16: 1 sign | 8 exponent | 7 mantissa (16 bits)
The exponent sets the range: the smallest and largest numbers the format can express. The mantissa sets the precision: how finely you can subdivide the gap between consecutive representable numbers.
What the exponent buys you
With 5 exponent bits, fp16’s normal range runs roughly 6e-5 to 6e4. Below
6e-5, you fall off into subnormals
or simply round to zero. fp16’s smallest subnormal is around 6e-8, so our
1e-9 gradient has nowhere to land at all: it becomes exactly 0.0. With 8
exponent bits, bf16’s normal range runs from about 1e-38 to 3e38 — same as
fp32. In bf16 the same gradient is stored as a slightly-rounded 1e-9, off by
under a percent, but there.
That gap matters because of how training works. The gradient of a deep
network is a product of many small numbers — one per layer. Multiply 50 small
numbers together and you get values like our 1e-9. In fp16, those tail
gradients silently round to zero, and any weight that depended on them
stops learning — not “learns slowly,” stops, because zero times any learning
rate is still zero. The standard fp16 fix is
loss scaling:
multiply the loss by some large constant (1024, 65536) so the resulting
gradients land in fp16’s representable window, then divide back out before
the optimizer step. It works, but you have to dynamically tune the scale —
too high and you overflow to infinity, too low and you underflow. Frameworks
shipped mixed-precision tooling (torch.cuda.amp, NVIDIA’s APEX) partly to
manage it.
Scale our 1e-9 gradient by 65536 and it becomes 6.5e-5 — just barely inside
fp16’s normal range. That’s the whole trick, and also the whole fragility: the
constant has to be re-tuned as the gradient distribution shifts during
training.
bf16 has fp32’s exponent, so any normal fp32 value has a bf16 counterpart at the same magnitude — no loss scaling needed, and mixed-precision recipes get dramatically simpler. (The one asterisk: bf16’s coverage of the very smallest subnormal values isn’t identical to fp32’s, and some hardware flushes denormals to zero outright. That corner is far below where training gradients live, which is why it rarely comes up.)
What the mantissa costs you
bf16’s 7 mantissa bits space consecutive representable values about 1 part in
128 apart — roughly 0.8% — so rounding to the nearest one costs at most half
that, about 0.4%. Our 1e-9 gradient is stored as whichever of those ticks is
closest. That sounds bad. The reason it’s fine for training is more
interesting than the usual “neural networks are noise-tolerant” hand-wave.
Stochastic gradient descent is, fundamentally, an algorithm that takes noisy gradient estimates from a small batch and adds them to weights with a small learning rate. The gradient you compute from one batch already differs substantially from the true full-dataset gradient — that sampling noise is the dominant error term in practice, and it’s typically far larger than the sub-1% rounding error bf16 introduces. Adding a small relative error to a number that is already a noisy estimate doesn’t move the needle. (The relative sizes vary by model and batch size; I’m describing the shape of the argument, not a measured ratio.)
There’s a real cost — accumulators inside matrix multiplies still need fp32
to avoid losing precision in long sums, which is why tensor cores compute
bf16 × bf16 → fp32 and only round back at the end. And the optimizer state
(Adam moments)
is usually kept in fp32 because it accumulates across millions of steps and
small biases compound. But the weights and activations — the things that
make up the bulk of memory and bandwidth — sit happily in bf16.
Why fp16 came first anyway
That leaves one loose end: if bf16 fits training better, why was fp16 the format sitting on the shelf? Because fp16 wasn’t designed for ML. It was designed for graphics — pixel shaders, HDR color, normal maps — where the values you store are bounded (a color channel, a unit-length normal vector) and you care about visible precision across that bounded range. For graphics, 5 exponent bits is plenty and 10 mantissa bits is a real upgrade over 8-bit fixed-point.
It got reused for ML because the hardware existed. The mismatch — that ML gradients want exponent and graphics shaders wanted mantissa — only became obvious once people tried to train large models in it. bf16 is the version that was designed once someone actually thought about the workload.
The seam: bf16 isn’t always enough
bf16 is great for training and inference of dense models. It starts to fray in a few places:
- Numerically delicate operations. Some normalization schemes (LayerNorm’s variance computation, log-sum-exp in softmax) benefit from being computed in fp32 even when surrounding ops are bf16. Most frameworks do this implicitly.
- Long-running accumulators. As mentioned, Adam state usually stays fp32. So do gradient accumulators across micro-batches in pipeline parallel training.
- Below 16 bits. Once you go to fp8, the exponent budget gets tight again — fp8 comes in two flavors (E4M3 with 4 exponent bits, E5M2 with 5) precisely because there’s no good single answer at that width. The bf16 lesson — keep exponent — pushes you toward E5M2 for things that need range, and E4M3 for things where you’ve already controlled the range with scaling.
The exact dates on which individual labs flipped their default training format from fp16-plus-loss-scaling to bf16 aren’t public. What is public is the direction: by the early 2020s bf16 was already the dominant choice in published training recipes that ran on bf16-capable hardware.
You started with bf16 = sign + 8-bit exponent + 7-bit mantissa. What did the
1e-9 gradient add to that line? — + the exponent is the part training actually needs. The mantissa bits bf16 gave up were mostly encoding noise
that minibatch sampling had already drowned out; the exponent bits it kept
were the difference between a gradient that exists and a gradient that is
exactly zero.
Check yourself
Before you go — someone hands you a workload where the values are all comfortably between 0.01 and 100, and they need to tell 1.0000 apart from 1.0005. Should they reach for bf16 because “it’s the format that won”?
Answer
No. Their range is tiny — four orders of magnitude, well inside fp16’s ~6e-5
to 6e4 normal range — and their precision requirement (about 1 part in 2000)
is finer than bf16’s ~1-in-128 spacing. This is the graphics-shaped workload fp16 was
actually designed for. bf16 “won” for a specific workload — training, where
gradients span dozens of orders of magnitude and every value is already noisy —
not in general. Pick the format by which axis your data actually stresses.
And one more: if bf16 removes the need for loss scaling, why do frameworks still keep Adam’s optimizer state in fp32?
Answer
Loss scaling solves a range problem; optimizer state has a precision problem. The moments accumulate over millions of steps, and each step adds a tiny update to a much larger running value. With only 7 mantissa bits, an update small relative to the accumulator rounds away entirely — the same swallowing effect that non-associativity posts describe, repeated a million times. Range is fine there; resolution isn’t.
Famous related terms
- fp32 —
fp32 = 1 sign + 8 exponent + 23 mantissa— the “default” float for decades. Still the format of choice for optimizer state and numerically sensitive reductions. - fp16 —
fp16 = 1 sign + 5 exponent + 10 mantissa— IEEE half. More precise than bf16, narrower range. Lost the training fight; still useful in graphics and some inference settings. - fp8 (E4M3 / E5M2) —
fp8 ≈ bf16's logic taken one step further— two flavors because at 8 bits you really do have to pick range or precision, not both. - Mixed precision training —
mixed precision = compute in low precision + accumulate in high precision. The recipe that makes any of this actually train. - Loss scaling —
loss scaling = multiply loss by big constant + divide gradients back + dynamically tune the constant. The fp16 workaround that bf16 made unnecessary. - Tensor core —
tensor core = matrix-multiply unit + native low-precision input + fp32 accumulate. The hardware that made low-precision training cheap.
Going deeper
- Mixed Precision Training (Micikevicius et al., arXiv 1710.03740, October 2017; ICLR 2018) — read this to see exactly what problem loss scaling was invented for, which is the same as seeing what bf16 got rid of. (The author list spans NVIDIA and Baidu, not NVIDIA alone.)
- Google’s bfloat16: The secret to high performance on Cloud TPUs (Cloud blog) — the clearest first-party answer to “why did Google pick this exponent/mantissa split and not another one.”
- Higham, Accuracy and Stability of Numerical Algorithms — the rabbit hole
for readers who want the general theory of why “range” and “precision” are
separate axes rather than one dial. Worth noting on the way in: fp16 is
IEEE-standardized (
binary16, IEEE 754-2008, retained in 754-2019) and bf16 is not — it’s a de facto standard codified by hardware vendors.