Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why FP8 training is stable

FP8 has only 256 representable values. Training a frontier model in it sounds insane — and it almost is. Here's the trick that makes it work.

AI & ML intermediate Apr 30, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never thought about how a computer stores a number. The prose below fills in the seams the pictures skip.

1 · The problem

A scale that reads in whole kilograms records a month of progress as nothing.

whole kilograms, and nothing finer four weeks on that scale week 1week 2week 3week 4 −100 g−100 g−100 g−100 g reads the samereads the samereads the samereads the same 400 g of real progress recorded as none every change was smaller than one mark, so each one rounded to zero An 8-bit number has 256 marks, and that is the whole instrument.
The scale is the running example for the whole post. Halving the bits roughly doubles how much arithmetic a chip will do and halves the bytes it must move — but 256 values, spread across numbers that range over many orders of magnitude, is a brutally coarse instrument.

2 · Failure 1

The instrument isn’t wrong. It’s aimed at the wrong part of the number line.

aimed too low: everything big saturates what the format reaches your numbers, all off the end — they all become the same value aimed too high: everything small vanishes what the format reaches your numbers, crowded at the bottom — they all round to zero The fix: give every batch of numbers its own slider. Divide by one stored number on the way in, multiply by it on the way out. now only the shape of the values has to fit in 256 marks, not their size picking that slider from last step’s numbers is cheap, and goes wrong exactly when the numbers jump
Saturating at the top and rounding to zero at the bottom are the same failure seen from two ends: a fixed window aimed at the wrong place. A per-group scale factor slides the window to wherever the numbers actually are, and how that factor is chosen is where the practical risk lives.

3 · Failure 2

Sliding fixes where the window sits. It can’t change how wide it is.

eight bits to spend — on fine marks, or on reach. Never both. fine marks, short reach reaches to about ±448 used going forward through the model those values cluster, so resolution is what pays coarse marks, long reach reaches to about ±57,344 used for the learning signals coming back those can be tiny in one layer and large in the next No slider fixes a reach problem — so ship two formats and use each where it fits. not everyone splits them this way: one public frontier recipe uses the fine-marks format for all its matrix multiplies
A scale factor moves the window; it cannot widen it. The two 8-bit formats spend their bits differently — one on resolution, one on reach — because the values going forward and the signals coming back have genuinely different shapes.

4 · Failure 3

The maths is sane now, and the model still barely learns.

keeping the running total on the coarse dial the weight, here the update is this wide — far less than one mark so the sum lands on the mark it started from the step rounds to zero, every single step this is the bathroom scale, exactly keeping the running total on a notepad 2.480931 2.480994 2.481058 2.481121 2.481185 every tiny step actually lands the coarse dial still does the multiplying — it just never has to remember anything “How do you train in 256 values?” — you don’t. Only the multiplying happens there.
An update is often many orders of magnitude smaller than the weight it is updating, so on a coarse dial it disappears. The weights the optimiser actually updates are kept in higher precision, and the 8-bit format is only what the multiplications run in — everything that has to remember something happens in more bits.

5 · Where the fixed version still breaks

You cannot weigh a pile of letters and one suitcase on the same dial.

one slider for the whole batch the suitcase the letters one outlier forces the range up, and every letter now reads zero one slider per small tile suitcase the outlier only spoils precision for its own tile, not the whole batch
A single scale factor assumes one number can describe a whole group’s distribution, which fails the moment a few entries dwarf the rest. Scaling in small tiles confines an outlier’s damage to its own neighbourhood — and how small a group should have to share a dial is the part of this that is still moving.

6 · Keep this card

The whole thing on one index card.

8-bit training = two formats — fine marks, and long reach + a slider per group of numbers + a notepad in higher precision and the live question: how small should a group be? the formats and the notepad look settled; the granularity, and low-precision bookkeeping, are still moving
Picture to keep: the coarse dial never has to measure anything huge or anything tiny — you slide the whole thing until the numbers you care about land in the middle of its range, and you keep the running total on a notepad in full precision instead of on the dial.

Why it exists

Picture a bathroom scale with a coarse dial — it reads in whole kilograms and nothing finer. Stand on it every morning for a month while you’re genuinely losing 100 grams a week, and the number never moves. Not because nothing happened, but because every real change was smaller than one mark on the dial, so each one rounded to zero. Four hundred grams of progress, recorded as none.

That scale is the running example for this post, because it is exactly the failure mode of training a neural network in 8 bits. Keep it in mind — we’ll add marks to it, slide its range around, and eventually give it a notepad.

A trillion-parameter LLM is mostly matrix multiplies. Halve the number of bits each number takes, and you roughly double how many of those multiplies you can do per second on the same chip — and halve the memory traffic feeding them. That is the entire prize. Going from FP32 down to 16-bit (FP16 or BF16) gave the field one such doubling. Going to 8 bits is the next one.

The reason “just use 8 bits” sounds insane is that an 8-bit float has exactly 256 representable values — 256 marks on the dial, and that’s the whole instrument. That’s not a typo. Across a tensor with millions of entries spanning many orders of magnitude — gradients near zero, activation outliers shooting into the thousands — 256 buckets is brutal. Naively cast everything to FP8 and the loss curve diverges.

(One place the scale analogy breaks, and it’s the useful place: a scale’s marks are evenly spaced, but a float’s aren’t. Floating-point values bunch up near zero and spread out as magnitude grows, so FP8 gives you fine resolution on small numbers and coarse resolution on big ones. That uneven spacing is what makes the tricks below possible at all.)

So FP8 training sat in the “in theory yes, in practice no” bucket for years. NVIDIA’s Hopper architecture (H100, announced March 2022) was the first NVIDIA GPU architecture with native FP8 tensor cores — hardware that could do the arithmetic said nothing about whether a real training run would converge on it. By late 2024 there were public examples: DeepSeek-V3 (arXiv preprint, Dec 27 2024) trained with an FP8 mixed-precision framework, and reported relative loss error below 0.25% versus BF16 in validation experiments on smaller proxy models (their Appendix B), not as a measurement of the full 671B run.

The interesting question isn’t can you train in FP8 — it’s what makes it stable, given that 256 values per tensor really is the constraint.

Why it matters now

Halving the precision of the matmuls halves the bytes they move and roughly doubles the arithmetic rate the hardware will do them at. That is not the same as halving the cost of a training run — optimizer state, gradient communication, non-GEMM work, and the scaling bookkeeping all stay where they were — but on a frontier run the GEMMs are a large enough share that the saving is substantial in wall-clock terms. It also expands what fits at all: more parameters, longer context, bigger batches, on the same hardware.

It matters past the giant labs too. FP8 is a supported serving format across the major inference stacks, and the same dial-sliding tricks are what keep a deployed model honest at roughly half the memory of BF16. Understanding why FP8 training works tells you why FP8 inference works, since they share the failure modes — the difference is that inference never has to accumulate an update, which is why it was solved first.

The short answer

FP8 training = (E4M3 + E5M2) + per-tensor scale factors + a higher-precision master copy of weights

Picture to keep: the coarse dial never has to measure anything big or anything tiny — you slide the whole thing until the numbers you care about land in the middle of its range, and you keep the running total on a notepad in full precision instead of on the dial.

Two FP8 formats, not one: E4M3 for the forward pass where precision matters, and E5M2 for gradients, which span a huge dynamic range. (DeepSeek-V3 deviates here and uses E4M3 for all of its FP8 GEMMs — more on that below.) Each tensor gets its own scale factor that slides its values into FP8’s narrow window before quantization, then slides back out. And the “real” weights — the ones the optimizer updates — live in higher precision; FP8 is just the format the matmuls run in. (How much higher varies: the original FP16 mixed-precision recipe and DeepSeek-V3 both keep FP32 master weights; BF16 is also used.)

How it works

Start from the naive version — cast every tensor to FP8, run the same training loop — and watch it fail three separate ways. Each fix is one term in the compression line, and each one is the scale analogy in a different costume.

Failure 1: the dial is pointed at the wrong range.

If your tensor’s largest absolute value is 800 and E4M3’s ceiling is 448, every value above 448 saturates to the same number — you’ve thrown away the tail. If the largest value is 0.001, almost everything rounds to zero — you’ve thrown away the body. This is trying to weigh a letter on a bathroom scale, or a person on a kitchen scale: the instrument is fine, it’s just aimed at the wrong part of the number line.

The fix: give every tensor its own scale factor. A single FP32 number s per tensor: store x / s in FP8, multiply by s on the way out. Now you only need the shape of the distribution to fit in 256 buckets, not its absolute magnitude. The dial slides to wherever the numbers actually are.

Picking s is the whole game. Derive it from the current tensor’s max-abs and you pay a reduction over the whole tensor every step. A common production trick is delayed scaling: keep a short history of recent max-abs values and derive s from that. Cheap, but structurally vulnerable in an obvious way: the scale you apply this step is derived from what the tensor looked like last step, so an unusually large activation makes the stored history wrong exactly when it matters. NVIDIA documents delayed scaling and current (just-in-time) scaling as the two supported options, which is the tell that the trade-off is real. No public post-mortem quantifies how often the delayed variant actually breaks a frontier run — treat the mechanism as sound and the frequency as unknown.

Failure 2: one dial can’t cover both jobs.

Scaling fixes where the range sits, not how wide it is. Forward activations are relatively well-behaved: they cluster, so you want fine marks. Gradients don’t — they can be vanishingly small in one layer and large in the next, within the same step. A format tuned for fine resolution simply doesn’t reach far enough for them, and no choice of s fixes a range problem.

The fix: two formats, and use each where it fits. FP8 has 8 bits to spend. You can spend more on the exponent (range) or more on the mantissa (precision); you can’t have both. NVIDIA standardized two splits:

The “hybrid” recipe — E4M3 for forward activations and weights, E5M2 for backward gradients — is the default in NVIDIA’s Transformer Engine. Forward values cluster in a manageable range; gradients can be tiny one layer and large the next, so they need the headroom.

Failure 3: the updates are smaller than one mark.

Now the matmuls are numerically sane and the loss stops exploding — and the model still barely learns. This is the bathroom scale from the opening, exactly. A gradient update is often many orders of magnitude smaller than the weight it’s updating. Add it to an FP8 weight and the sum lands on the same representable value it started from. The step rounds to zero, every step, and a month of training records no progress.

The fix: keep the running total off the dial. The matmuls run in FP8. The weights themselves — the parameters the optimizer updates — are kept in BF16 or FP32. Each step you cast down to FP8 to compute, then apply the gradient update to the high-precision master copy. Optimizer state (Adam’s moments) also stays in higher precision. FP8 buys you the matmul throughput; the master copy buys you the ability to actually accumulate learning. This is the same pattern as FP16 mixed-precision training (Micikevicius et al., 2017) — FP8 just pushes it further.

That’s the whole recipe, and it’s why “how do you train stably in 256 values?” has a boring answer: you don’t. Only the multiplications happen in 256 values. Everything that has to remember something happens in more.

The seam: where the fixed version still breaks.

Per-tensor scaling assumes one number can describe a whole tensor’s distribution. That’s wrong when a tensor has outliers — a few entries far larger than the rest. Back to the scale: you can’t weigh a pile of letters and one suitcase on the same dial. The suitcase forces the range up, and every letter now reads zero. Activation outliers in transformers are a well-documented version of this problem.

DeepSeek-V3’s contribution was to scale at finer granularity: 1×128 tiles for activations, 128×128 blocks for weights. Each tile/block gets its own scale, so an outlier only ruins precision for its neighborhood, not the whole tensor. They also use E4M3 (not the hybrid E4M3/E5M2 split) for all FP8 GEMMs, arguing the finer-grained scaling reclaimed enough effective range; some non-GEMM ops still stay in BF16/FP32 in their framework. Whether this fine-grained approach is now the dominant production recipe across other frontier labs, I don’t have public information to say — labs are quiet about training stacks.

You started with FP8 training = (E4M3 + E5M2) + per-tensor scale factors + a higher-precision master copy of weights. What did the three failures add to that line? — + scaling granularity, the term that’s still moving fastest. The formats and the master copy are the settled-looking parts; the live engineering question is how small a group of numbers should have to share a dial, and delayed vs. current scaling, per-tile, and per-block are all answers to that one question. (Settled-looking is the honest word: low-precision optimizer state and FP4/NVFP4 recipes are both active enough that I wouldn’t call any of this closed.)

Check yourself

Before you go — someone proposes storing the optimizer state in FP8 too, since it’s a big chunk of training memory. Which of the three failures does that walk straight back into?

Answer

Failure 3, the bathroom scale. Adam’s moments are running accumulators: their whole job is to integrate a long series of small contributions. Round each contribution to the nearest of 256 marks and the small ones vanish, exactly as the master weights would. That’s why the standard recipe keeps optimizer state in higher precision even though it’s memory-expensive. (This doesn’t mean low-precision optimizer state is impossible — people work on it — but it needs its own machinery, not just a cast.)

And a diagnostic one: your FP8 run is stable for 10,000 steps, then the loss spikes to NaN in a single step and never recovers. You’ve already ruled out a bad data batch and an optimizer bug. What’s the FP8-specific suspect, and why does the shape of the failure point at it?

Answer

The scale factors, and delayed scaling in particular. The shape is the clue: a run that was stable for a long time and then broke in one step doesn’t look like a gradual precision problem, it looks like a range problem that arrived suddenly. If s comes from a history of recent max-abs values, one iteration with an unusually large activation leaves the stored history wrong for the next step, and values overflow the format’s ceiling. Checking whether the spike coincides with an outlier in the previous step’s amax history is the cheap first test. (Sudden NaNs have plenty of non-FP8 causes too — this is the FP8-flavoured hypothesis, not the only one.)

Going deeper