Why FP8 training is stable
FP8 has only 256 representable values. Training a frontier model in it sounds insane — and it almost is. Here's the trick that makes it work.
On this page
The picture version
Six pictures for a reader who has never thought about how a computer stores a number. The prose below fills in the seams the pictures skip.
1 · The problem
A scale that reads in whole kilograms records a month of progress as nothing.
2 · Failure 1
The instrument isn’t wrong. It’s aimed at the wrong part of the number line.
3 · Failure 2
Sliding fixes where the window sits. It can’t change how wide it is.
4 · Failure 3
The maths is sane now, and the model still barely learns.
5 · Where the fixed version still breaks
You cannot weigh a pile of letters and one suitcase on the same dial.
6 · Keep this card
The whole thing on one index card.
Why it exists
Picture a bathroom scale with a coarse dial — it reads in whole kilograms and nothing finer. Stand on it every morning for a month while you’re genuinely losing 100 grams a week, and the number never moves. Not because nothing happened, but because every real change was smaller than one mark on the dial, so each one rounded to zero. Four hundred grams of progress, recorded as none.
That scale is the running example for this post, because it is exactly the failure mode of training a neural network in 8 bits. Keep it in mind — we’ll add marks to it, slide its range around, and eventually give it a notepad.
A trillion-parameter LLM is mostly matrix multiplies. Halve the number of bits each number takes, and you roughly double how many of those multiplies you can do per second on the same chip — and halve the memory traffic feeding them. That is the entire prize. Going from FP32 down to 16-bit (FP16 or BF16) gave the field one such doubling. Going to 8 bits is the next one.
The reason “just use 8 bits” sounds insane is that an 8-bit float has exactly 256 representable values — 256 marks on the dial, and that’s the whole instrument. That’s not a typo. Across a tensor with millions of entries spanning many orders of magnitude — gradients near zero, activation outliers shooting into the thousands — 256 buckets is brutal. Naively cast everything to FP8 and the loss curve diverges.
(One place the scale analogy breaks, and it’s the useful place: a scale’s marks are evenly spaced, but a float’s aren’t. Floating-point values bunch up near zero and spread out as magnitude grows, so FP8 gives you fine resolution on small numbers and coarse resolution on big ones. That uneven spacing is what makes the tricks below possible at all.)
So FP8 training sat in the “in theory yes, in practice no” bucket for years. NVIDIA’s Hopper architecture (H100, announced March 2022) was the first NVIDIA GPU architecture with native FP8 tensor cores — hardware that could do the arithmetic said nothing about whether a real training run would converge on it. By late 2024 there were public examples: DeepSeek-V3 (arXiv preprint, Dec 27 2024) trained with an FP8 mixed-precision framework, and reported relative loss error below 0.25% versus BF16 in validation experiments on smaller proxy models (their Appendix B), not as a measurement of the full 671B run.
The interesting question isn’t can you train in FP8 — it’s what makes it stable, given that 256 values per tensor really is the constraint.
Why it matters now
Halving the precision of the matmuls halves the bytes they move and roughly doubles the arithmetic rate the hardware will do them at. That is not the same as halving the cost of a training run — optimizer state, gradient communication, non-GEMM work, and the scaling bookkeeping all stay where they were — but on a frontier run the GEMMs are a large enough share that the saving is substantial in wall-clock terms. It also expands what fits at all: more parameters, longer context, bigger batches, on the same hardware.
It matters past the giant labs too. FP8 is a supported serving format across the major inference stacks, and the same dial-sliding tricks are what keep a deployed model honest at roughly half the memory of BF16. Understanding why FP8 training works tells you why FP8 inference works, since they share the failure modes — the difference is that inference never has to accumulate an update, which is why it was solved first.
The short answer
FP8 training = (E4M3 + E5M2) + per-tensor scale factors + a higher-precision master copy of weights
Picture to keep: the coarse dial never has to measure anything big or anything tiny — you slide the whole thing until the numbers you care about land in the middle of its range, and you keep the running total on a notepad in full precision instead of on the dial.
Two FP8 formats, not one: E4M3 for the forward pass where precision matters, and E5M2 for gradients, which span a huge dynamic range. (DeepSeek-V3 deviates here and uses E4M3 for all of its FP8 GEMMs — more on that below.) Each tensor gets its own scale factor that slides its values into FP8’s narrow window before quantization, then slides back out. And the “real” weights — the ones the optimizer updates — live in higher precision; FP8 is just the format the matmuls run in. (How much higher varies: the original FP16 mixed-precision recipe and DeepSeek-V3 both keep FP32 master weights; BF16 is also used.)
How it works
Start from the naive version — cast every tensor to FP8, run the same training loop — and watch it fail three separate ways. Each fix is one term in the compression line, and each one is the scale analogy in a different costume.
Failure 1: the dial is pointed at the wrong range.
If your tensor’s largest absolute value is 800 and E4M3’s ceiling is 448, every value above 448 saturates to the same number — you’ve thrown away the tail. If the largest value is 0.001, almost everything rounds to zero — you’ve thrown away the body. This is trying to weigh a letter on a bathroom scale, or a person on a kitchen scale: the instrument is fine, it’s just aimed at the wrong part of the number line.
The fix: give every tensor its own scale factor. A single FP32 number s per tensor: store x / s in FP8, multiply by s on the way out. Now you only need the shape of the distribution to fit in 256 buckets, not its absolute magnitude. The dial slides to wherever the numbers actually are.
Picking s is the whole game. Derive it from the current tensor’s max-abs and you pay a reduction over the whole tensor every step. A common production trick is delayed scaling: keep a short history of recent max-abs values and derive s from that. Cheap, but structurally vulnerable in an obvious way: the scale you apply this step is derived from what the tensor looked like last step, so an unusually large activation makes the stored history wrong exactly when it matters. NVIDIA documents delayed scaling and current (just-in-time) scaling as the two supported options, which is the tell that the trade-off is real. No public post-mortem quantifies how often the delayed variant actually breaks a frontier run — treat the mechanism as sound and the frequency as unknown.
Failure 2: one dial can’t cover both jobs.
Scaling fixes where the range sits, not how wide it is. Forward activations are relatively well-behaved: they cluster, so you want fine marks. Gradients don’t — they can be vanishingly small in one layer and large in the next, within the same step. A format tuned for fine resolution simply doesn’t reach far enough for them, and no choice of s fixes a range problem.
The fix: two formats, and use each where it fits. FP8 has 8 bits to spend. You can spend more on the exponent (range) or more on the mantissa (precision); you can’t have both. NVIDIA standardized two splits:
- E4M3 — 4 exponent bits, 3 mantissa bits. Range up to about ±448. More precision, less range.
- E5M2 — 5 exponent bits, 2 mantissa bits. Range up to about ±57,344. More range, less precision.
The “hybrid” recipe — E4M3 for forward activations and weights, E5M2 for backward gradients — is the default in NVIDIA’s Transformer Engine. Forward values cluster in a manageable range; gradients can be tiny one layer and large the next, so they need the headroom.
Failure 3: the updates are smaller than one mark.
Now the matmuls are numerically sane and the loss stops exploding — and the model still barely learns. This is the bathroom scale from the opening, exactly. A gradient update is often many orders of magnitude smaller than the weight it’s updating. Add it to an FP8 weight and the sum lands on the same representable value it started from. The step rounds to zero, every step, and a month of training records no progress.
The fix: keep the running total off the dial. The matmuls run in FP8. The weights themselves — the parameters the optimizer updates — are kept in BF16 or FP32. Each step you cast down to FP8 to compute, then apply the gradient update to the high-precision master copy. Optimizer state (Adam’s moments) also stays in higher precision. FP8 buys you the matmul throughput; the master copy buys you the ability to actually accumulate learning. This is the same pattern as FP16 mixed-precision training (Micikevicius et al., 2017) — FP8 just pushes it further.
That’s the whole recipe, and it’s why “how do you train stably in 256 values?” has a boring answer: you don’t. Only the multiplications happen in 256 values. Everything that has to remember something happens in more.
The seam: where the fixed version still breaks.
Per-tensor scaling assumes one number can describe a whole tensor’s distribution. That’s wrong when a tensor has outliers — a few entries far larger than the rest. Back to the scale: you can’t weigh a pile of letters and one suitcase on the same dial. The suitcase forces the range up, and every letter now reads zero. Activation outliers in transformers are a well-documented version of this problem.
DeepSeek-V3’s contribution was to scale at finer granularity: 1×128 tiles for activations, 128×128 blocks for weights. Each tile/block gets its own scale, so an outlier only ruins precision for its neighborhood, not the whole tensor. They also use E4M3 (not the hybrid E4M3/E5M2 split) for all FP8 GEMMs, arguing the finer-grained scaling reclaimed enough effective range; some non-GEMM ops still stay in BF16/FP32 in their framework. Whether this fine-grained approach is now the dominant production recipe across other frontier labs, I don’t have public information to say — labs are quiet about training stacks.
You started with FP8 training = (E4M3 + E5M2) + per-tensor scale factors + a higher-precision master copy of weights. What did the three failures add to that line? — + scaling granularity, the term that’s still moving fastest. The formats and the master copy are the settled-looking parts; the live engineering question is how small a group of numbers should have to share a dial, and delayed vs. current scaling, per-tile, and per-block are all answers to that one question. (Settled-looking is the honest word: low-precision optimizer state and FP4/NVFP4 recipes are both active enough that I wouldn’t call any of this closed.)
Check yourself
Before you go — someone proposes storing the optimizer state in FP8 too, since it’s a big chunk of training memory. Which of the three failures does that walk straight back into?
Answer
Failure 3, the bathroom scale. Adam’s moments are running accumulators: their whole job is to integrate a long series of small contributions. Round each contribution to the nearest of 256 marks and the small ones vanish, exactly as the master weights would. That’s why the standard recipe keeps optimizer state in higher precision even though it’s memory-expensive. (This doesn’t mean low-precision optimizer state is impossible — people work on it — but it needs its own machinery, not just a cast.)
And a diagnostic one: your FP8 run is stable for 10,000 steps, then the loss spikes to NaN in a single step and never recovers. You’ve already ruled out a bad data batch and an optimizer bug. What’s the FP8-specific suspect, and why does the shape of the failure point at it?
Answer
The scale factors, and delayed scaling in particular. The shape is the clue: a run that was stable for a long time and then broke in one step doesn’t look like a gradual precision problem, it looks like a range problem that arrived suddenly. If s comes from a history of recent max-abs values, one iteration with an unusually large activation leaves the stored history wrong for the next step, and values overflow the format’s ceiling. Checking whether the spike coincides with an outlier in the previous step’s amax history is the cheap first test. (Sudden NaNs have plenty of non-FP8 causes too — this is the FP8-flavoured hypothesis, not the only one.)
Famous related terms
- Mixed precision training —
mixed precision = low-precision matmul + high-precision master weights— the general pattern; FP8 is its current frontier. - Loss scaling —
loss scaling = multiply loss by k + divide gradients by k— the FP16-era trick to keep tiny gradients from underflowing to zero. Per-tensor scaling generalizes the same idea per-tensor instead of globally. - Quantization —
quantization = high-precision values + scale + low-precision storage— see why quantization works. Inference quantization and FP8 training stability are the same numerical problem viewed from two ends. - Transformer Engine —
Transformer Engine ≈ NVIDIA library + drop-in FP8 layers + automatic scale tracking. Hides the bookkeeping behind drop-in layers; the reason most teams don’t implement FP8 from scratch. - FP4 —
FP4 = 4-bit float + scale factor + even-narrower window. 4-bit floats, supported on Blackwell. Same playbook, less margin. Whether FP4 training (not just inference) is broadly stable is still an active question as of early 2026, with no definitive public answer yet.
Going deeper
- DeepSeek-V3 Technical Report — the primary source for the fine-grained scaling story, and the best public answer to “what does a real frontier-scale FP8 recipe actually look like.”
- NVIDIA Transformer Engine FP8 primer — answers “what do I actually have to configure,” walking through the hybrid formats and delayed scaling as code rather than theory.
- Rabbit hole: To FP8 and Back Again (Lee et al., 2024) — answers “how does FP8 training actually fail when it fails,” which is the question the marketing material never covers.