Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why bf16 won the training format wars

Half-precision floats came in two flavors: fp16, which had been around for years, and bf16, which kept fp32's exponent and threw away mantissa bits. The less-precise format won. Here's why that's not a typo.

Computer Science intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never thought about how many digits a number gets. The prose below fills in the seams the pictures skip.

1 · The problem

Same sixteen bits, split two ways. One of them loses your number entirely.

a gradient of 1e−9 stored in one 16-bit format 0.0 stored in the other 1e−9 Neither format is bigger. They spend the same sixteen bits differently. and a gradient that rounds to zero contributes nothing — the model simply stops learning from it
Two formats of identical size, and a number that survives in one and vanishes in the other. That 1e-9 gradient is the running example, and the whole difference comes from how the bits are divided.

2 · What the bits are for

Some bits say how big. The rest say how precisely.

sixteen bits, two ways of dividing them format A ± 5 bits: how big 10 bits: how precisely format B ± 8 bits: how big 7 bits: how precisely Every bit you give to reach is a bit you take from resolution. There is no third pot to take from.
One bit records the sign; the rest are split between how large a number can get and how finely it can be pinned down. The two formats make opposite calls on that split, and everything downstream follows from it.

3 · What reach buys

Two rulers with the same number of ticks. One is short.

format A reaches this far format B reaches this far 1e−9 lives here Off the end of one ruler entirely. Comfortably inside the other.
Spending bits on reach stretches the range of magnitudes a format can represent at all. A number outside that range doesn’t get stored roughly — it gets stored as zero, and a zero gradient teaches the model nothing.

4 · What reach costs

Wider ruler, coarser ticks. About three decimal digits, and that’s your lot.

fine ticks: values close together stay distinct coarse ticks: nearby values collapse onto the same one A coarse tick is still a tick. A number off the ruler is nothing at all.
The wider format pays for its reach in resolution — roughly three decimal digits rather than four. That trade is worth it when the alternative is losing the number completely, which is the judgement the whole comparison turns on.

5 · Why the narrow one came first

It was designed for pictures, where nothing is ever that small.

what it was built for colours and coordinates all in a narrow band of sizes precision matters, reach doesn’t what it got used for later training signals spanning many orders of magnitude reach is exactly what runs out The format wasn’t wrong. It was answering a different question.
The narrower format predates the workload that broke it, and for graphics its choice was the right one. The mismatch appeared when the same sixteen bits were pointed at numbers spanning a far wider range — which is why practitioners had to bolt on tricks to keep values inside the ruler.

6 · Keep this card

The whole thing on one index card.

the swap = same 16 bits, but more for reach — and correspondingly fewer for resolution ∴ because a coarse number beats a lost one And it isn’t free: reach alone doesn’t save you if the values drift far enough. which is why 8-bit formats went back to scaling values into range rather than widening the ruler again
Picture to keep: two rulers with the same number of tick marks — one short and finely etched, the other stretching from the atomic to the astronomical with coarse ticks. The 1e-9 gradient falls off the end of the first and lands in the middle of the second, and a coarse tick is still a tick. Where it breaks: a float’s ticks aren’t evenly spaced, they widen with magnitude — which is why big numbers swallow small ones.

Why it exists

Punch 0.0000001 × 0.0000001 into your phone’s calculator. It doesn’t print a long crawl of zeros and it doesn’t give up — it shows something like 1e-14. The display has a fixed number of character slots, and scientific notation spends a few of them on the exponent so the rest can carry digits. You just traded precision for range without thinking about it, and you got a number you could still read.

That’s the only decision a 16-bit float format ever makes: given a fixed bit budget, how much goes to how big or small a number you can write down and how much to how precisely you can write it. fp16 spent more on precision. bf16 spent more on range. For training neural networks, range turned out to matter far more — which is why the less precise format won.

Keep one number in your pocket for the rest of this post: a single weight gradient that comes out to 1e-9. Everything below is the story of whether that number survives the trip from the backward pass into the optimizer.

Here’s the puzzle. You’re training a neural network. You’d like to use 16-bit floats instead of 32-bit floats — half the memory, half the bandwidth, often more than double the throughput on modern matrix engines. There are two candidates sitting on the shelf:

Read those numbers carefully. bf16 has fewer mantissa bits than fp16 — 7 vs. 10. It’s the less precise of the two. And yet bf16 is now the default mixed-precision format on every major training stack I’ve seen public recipes for, and fp16 is treated as the format you have to work around.

That should read as backwards, because the entire historical case for floating point was “trade range for precision smartly, but get plenty of both.” bf16 deliberately throws precision away to keep range. Why was that the right trade?

The short version: training a neural network is a lot less sensitive to how precisely you can represent a number than to how small or large a number you can represent at all. fp16’s narrower exponent squeezes from both ends — small gradients underflow to zero, and scaled-up values can overflow to infinity — so engineers spent years inventing workarounds (loss scaling, mixed-precision recipes) to keep training inside the window. bf16 just gives the gradients somewhere to live, accepts that the last couple of bits of the mantissa are noise anyway, and gets out of the way.

Why it matters now

If you train models, this is one of the few hardware decisions that’s still visible in your code. PyTorch’s torch.bfloat16 and torch.float16 are not interchangeable, and picking the wrong one is the difference between a stable run and one that quietly diverges to NaN partway through.

It also explains why specific accelerators became popular when they did. Google developed bfloat16 for Cloud TPU v2 and v3; Cloud TPU went publicly available in beta in February 2018. NVIDIA’s tensor cores gained native bf16 with the A100 (Ampere, 2020). After that point, “use mixed precision” stopped meaning “fp16 with a bag of tricks” and started meaning “bf16, which Just Works.” Frameworks shifted defaults, and by the H100 generation (2022) bf16 was the common choice in publicly described large-model training recipes running on bf16-capable hardware.

It matters even outside training. Inference is moving toward fp8 and below (fp8 in H100, fp4 announced for Blackwell), and the design decisions there inherit bf16’s lesson: when you have to throw bits away, throw away precision before you throw away range.

The short answer

bf16 = sign + 8-bit exponent (same as fp32) + 7-bit mantissa

Picture to keep: two rulers with the same number of tick marks. fp16’s ruler is short but finely etched; bf16’s stretches from the atomic to the astronomical with coarse ticks. Our 1e-9 gradient falls off the end of fp16’s ruler entirely, and lands comfortably in the middle of bf16’s — a coarse tick is still a tick. (Where the ruler breaks: a float’s ticks aren’t evenly spaced like a real ruler’s. They’re spaced proportionally, so the gaps grow with the magnitude — which is why big numbers swallow small ones.)

bf16 is essentially fp32 with the bottom 16 bits of the mantissa lopped off (plus a rounding rule, and on some hardware a flush-to-zero treatment of denormals). It has the same exponent range as fp32 — the same ability to represent very small gradients and very large activations — but only enough mantissa precision for roughly 2–3 significant decimal digits. That’s plenty for gradient descent, which spends its life adding noisy small updates to noisy current weights. It is not plenty for, say, a stiff differential equation solver — but neural network training isn’t that.

How it works

The oddity to keep in view is that the less precise format won, and the reason is that the exponent and the mantissa fail in completely different ways. A floating-point number is sign × mantissa × 2^exponent. You get a fixed bit budget; you split it between mantissa and exponent.

fp32:  1 sign | 8 exponent | 23 mantissa     (32 bits total)
fp16:  1 sign | 5 exponent | 10 mantissa     (16 bits)
bf16:  1 sign | 8 exponent |  7 mantissa     (16 bits)

The exponent sets the range: the smallest and largest numbers the format can express. The mantissa sets the precision: how finely you can subdivide the gap between consecutive representable numbers.

What the exponent buys you

With 5 exponent bits, fp16’s normal range runs roughly 6e-5 to 6e4. Below 6e-5, you fall off into subnormals or simply round to zero. fp16’s smallest subnormal is around 6e-8, so our 1e-9 gradient has nowhere to land at all: it becomes exactly 0.0. With 8 exponent bits, bf16’s normal range runs from about 1e-38 to 3e38 — same as fp32. In bf16 the same gradient is stored as a slightly-rounded 1e-9, off by under a percent, but there.

That gap matters because of how training works. The gradient of a deep network is a product of many small numbers — one per layer. Multiply 50 small numbers together and you get values like our 1e-9. In fp16, those tail gradients silently round to zero, and any weight that depended on them stops learning — not “learns slowly,” stops, because zero times any learning rate is still zero. The standard fp16 fix is loss scaling: multiply the loss by some large constant (1024, 65536) so the resulting gradients land in fp16’s representable window, then divide back out before the optimizer step. It works, but you have to dynamically tune the scale — too high and you overflow to infinity, too low and you underflow. Frameworks shipped mixed-precision tooling (torch.cuda.amp, NVIDIA’s APEX) partly to manage it. Scale our 1e-9 gradient by 65536 and it becomes 6.5e-5 — just barely inside fp16’s normal range. That’s the whole trick, and also the whole fragility: the constant has to be re-tuned as the gradient distribution shifts during training.

bf16 has fp32’s exponent, so any normal fp32 value has a bf16 counterpart at the same magnitude — no loss scaling needed, and mixed-precision recipes get dramatically simpler. (The one asterisk: bf16’s coverage of the very smallest subnormal values isn’t identical to fp32’s, and some hardware flushes denormals to zero outright. That corner is far below where training gradients live, which is why it rarely comes up.)

What the mantissa costs you

bf16’s 7 mantissa bits space consecutive representable values about 1 part in 128 apart — roughly 0.8% — so rounding to the nearest one costs at most half that, about 0.4%. Our 1e-9 gradient is stored as whichever of those ticks is closest. That sounds bad. The reason it’s fine for training is more interesting than the usual “neural networks are noise-tolerant” hand-wave.

Stochastic gradient descent is, fundamentally, an algorithm that takes noisy gradient estimates from a small batch and adds them to weights with a small learning rate. The gradient you compute from one batch already differs substantially from the true full-dataset gradient — that sampling noise is the dominant error term in practice, and it’s typically far larger than the sub-1% rounding error bf16 introduces. Adding a small relative error to a number that is already a noisy estimate doesn’t move the needle. (The relative sizes vary by model and batch size; I’m describing the shape of the argument, not a measured ratio.)

There’s a real cost — accumulators inside matrix multiplies still need fp32 to avoid losing precision in long sums, which is why tensor cores compute bf16 × bf16 → fp32 and only round back at the end. And the optimizer state (Adam moments) is usually kept in fp32 because it accumulates across millions of steps and small biases compound. But the weights and activations — the things that make up the bulk of memory and bandwidth — sit happily in bf16.

Why fp16 came first anyway

That leaves one loose end: if bf16 fits training better, why was fp16 the format sitting on the shelf? Because fp16 wasn’t designed for ML. It was designed for graphics — pixel shaders, HDR color, normal maps — where the values you store are bounded (a color channel, a unit-length normal vector) and you care about visible precision across that bounded range. For graphics, 5 exponent bits is plenty and 10 mantissa bits is a real upgrade over 8-bit fixed-point.

It got reused for ML because the hardware existed. The mismatch — that ML gradients want exponent and graphics shaders wanted mantissa — only became obvious once people tried to train large models in it. bf16 is the version that was designed once someone actually thought about the workload.

The seam: bf16 isn’t always enough

bf16 is great for training and inference of dense models. It starts to fray in a few places:

The exact dates on which individual labs flipped their default training format from fp16-plus-loss-scaling to bf16 aren’t public. What is public is the direction: by the early 2020s bf16 was already the dominant choice in published training recipes that ran on bf16-capable hardware.

You started with bf16 = sign + 8-bit exponent + 7-bit mantissa. What did the 1e-9 gradient add to that line? — + the exponent is the part training actually needs. The mantissa bits bf16 gave up were mostly encoding noise that minibatch sampling had already drowned out; the exponent bits it kept were the difference between a gradient that exists and a gradient that is exactly zero.

Check yourself

Before you go — someone hands you a workload where the values are all comfortably between 0.01 and 100, and they need to tell 1.0000 apart from 1.0005. Should they reach for bf16 because “it’s the format that won”?

Answer

No. Their range is tiny — four orders of magnitude, well inside fp16’s ~6e-5 to 6e4 normal range — and their precision requirement (about 1 part in 2000) is finer than bf16’s ~1-in-128 spacing. This is the graphics-shaped workload fp16 was actually designed for. bf16 “won” for a specific workload — training, where gradients span dozens of orders of magnitude and every value is already noisy — not in general. Pick the format by which axis your data actually stresses.

And one more: if bf16 removes the need for loss scaling, why do frameworks still keep Adam’s optimizer state in fp32?

Answer

Loss scaling solves a range problem; optimizer state has a precision problem. The moments accumulate over millions of steps, and each step adds a tiny update to a much larger running value. With only 7 mantissa bits, an update small relative to the accumulator rounds away entirely — the same swallowing effect that non-associativity posts describe, repeated a million times. Range is fine there; resolution isn’t.

Going deeper