Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why quantization works

Stuffing a 70-billion-parameter model into 4-bit weights sounds like it should ruin it. It mostly doesn't — and the reason is more about how the model gets used at inference than about the math of rounding.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has only ever seen the file-size column. The prose below fills in the seams the pictures skip.

1 · The problem

Cut every number to a quarter of its precision. Why isn’t it ruined?

the download page model-70B-f16 140 GB model-70B-Q8 ~70 GB model-70B-Q4 ~35 GB and the comments underneath say it’s basically fine do that to anything else a photograph visibly falls apart audio hisses your GPU has 24 GB, so this is the whole question Why doesn’t the model just become gibberish? the answer is an asymmetry: the cost of quantization is small and the benefit is large — and both halves are specific to how this kind of model gets used
Not a universal law about neural networks. Two separate facts have to hold at once: a trained model tolerates weight noise far better than its bit-width suggests, and inference is bottlenecked on exactly the thing quantization relieves.

2 · Why it survives at all

Each error is tiny and there are billions of them.

one output value is a sum over thousands of terms w₁x₁ w₂x₂ w₃x₃ w₄x₄ … each one now carries a small rounding error, up or down summed most of the noise cancels The central limit theorem is the shape of this argument, not a proof of it. the errors aren’t independent and the model is learned, so the maths isn’t strict and there is no clean theory saying how many bits a model of size N can lose people quantize, measure the benchmark drop, and publish tables. the envelope is measured, not derived.
The older literature calls this “over-parameterization buys robustness” — the same fact pruning and distillation lean on. Networks tolerate weight noise far better than the bit-level maths would predict, and nobody has explained precisely how much better.

3 · First repair

Pick the range per small group, not once for the whole model.

one range for everything one outlier buckets have to span the whole range, so every ordinary weight rounds coarsely a tight range per group each group gets its own scale factor, so a wide channel can’t drag down the rest This works because trained weights cluster tightly around zero. the same reason JPEG works on photos: the signal has structure, so you can spend bits where they matter where that analogy breaks: JPEG drops detail your eye can’t resolve. nothing tells you in advance what the model can’t resolve.
Which bits are safe to drop is settled by measuring, not by theory — quantization has no perceptual model to appeal to. If weights were spread uniformly over a huge range this scheme would be hopeless; the friendly bell-shaped distribution is what you are exploiting.

4 · Why you’d bother

The GPU was waiting on memory, not on maths.

one token, batch of one, dense 70B model memory holds the weights compute units mostly idle 140 GB streamed, every single token at ~3.35 TB/s that is ~40 ms per token even if the maths took zero time so store the weights in 4 bits and unpack them inside the kernel 4× less to read you pay extra compute to dequantize back to 16-bit — but compute was sitting idle anyway, so the speedup tracks the bit reduction, minus whatever the kernel gives back in unpacking and non-weight traffic A compute-bound workload would get almost none of this.
That asymmetry is what makes quantization shine for LLM decode specifically. A small image classifier on a fast GPU wasn’t bottlenecked on reading its weights, so shrinking them buys little. In LLM decode you absolutely were, so the saving converts almost directly into throughput. Roofline figures are an order-of-magnitude sketch, not a benchmark.

5 · Where it actually breaks

A handful of dimensions refuse to be average.

activation dimensions a few specific dimensions, far off the scale of the rest quantize naively and the range must stretch so precision is wasted on the overwhelming majority quality collapses The fix: route the outliers through higher precision, quantize the rest. LLM.int8() reports keeping more than 99.9% of values in 8-bit while preserving accuracy up to 175B parameters later schemes respond to the same observation: SmoothQuant shifts difficulty between weights and activations, AWQ protects the most salient channels
This is the seam in the whole story. Quantization “just works” on the average weight; the engineering is almost entirely about the few percent where it doesn’t — which is also why aggressive schemes below 4 bits are still a research area rather than a default.

6 · Keep this card

The whole thing on one index card.

quantization = lower-precision storage of the weights + a calibrated map back to real numbers + a quality budget you spend carefully ∴ cheap because decode was memory-bound all along
Picture to keep: a wall of dimmer switches. Instead of sliding anywhere, each clicks into 16 detents — but the electrician sets each panel’s range to match how bright those lamps actually get, so almost nobody ends up at the wrong brightness. Where it breaks: a room with one wrong dimmer just looks slightly off, whereas a handful of specific weights and activations matter far more than the rest.

Why it exists

You go to download an open-weights model to run on your own machine. The full 70-billion-parameter release is around 140 GB of raw weights and your GPU has 24. Then you notice the file list has half a dozen other versions of the same model — Q4, Q5, Q8 — and the 4-bit one is roughly a quarter the size (a bit more, once you count the metadata each scheme carries), and the comments underneath say it’s basically fine. That’s a 4× cut in the precision of every single number in the model. If you did that to a photograph, it would visibly fall apart. If you did it to audio, it would hiss. Why doesn’t the model just become gibberish?

It’s a fair question, and the answer is the whole reason those Q4 files exist at all. The short version: a trained LLM turns out to be much more robust to weight noise than its precision suggests, and inference turns out to be bottlenecked on something quantization directly relieves — moving bytes, not crunching numbers. So the cost of quantization is small and the benefit is large, in a way that’s specific to how this kind of model gets used, not a universal law about neural networks.

This post is about why that asymmetry exists. The mechanics of any particular scheme (GPTQ, AWQ, bitsandbytes, FP8, INT4, MXFP4) you can look up; the load-bearing intuition is what makes them all keep working.

Why it matters now

Quantization is the lever that decides whether a given model actually runs on a given GPU. A 70B model in 16-bit weights is roughly 140 GB. An H100 SXM has 80 GB of VRAM (post). The model does not fit. In INT8 the weights are around 70 GB — technically under the line, but weights aren’t the only thing that has to live there, so in practice you have almost nothing left for KV cache or activations. In INT4 they’re around 35 GB and you have real headroom for users. The whole “can a hobbyist run this on one card?” question is decided here.

The same lever shows up in serving economics. Dense decode at small batch is memory-bandwidth bound — the GPU has to stream the model’s weights through the compute units to produce each token — so cutting weight bytes by 4× cuts the per-token weight traffic by up to 4×. Treat that as a ceiling, not a promise: scale metadata, non-weight traffic, and kernel quality all eat into it. Still, it’s why quantized builds are the norm rather than the exception in on-device runtimes (llama.cpp, MLX) and why serving stacks ship quantization support out of the box. NVIDIA’s Hopper generation even introduced FP8 in hardware so that training could move below 16-bit (NVIDIA H100 Transformer Engine announcement).

The short answer

quantization = lower-precision storage of weights + a calibrated map back to real numbers + a quality budget you spend carefully

Picture to keep: a wall of dimmer switches. Instead of a knob that slides anywhere, each one now clicks into 16 detents — but the electrician sets the range of each panel of dimmers to match how bright those lamps actually get, so almost nobody ends up at the wrong brightness. Where the analogy breaks: a room with one wrong dimmer just looks slightly off, whereas a handful of specific weights and activations matter far more than the rest — which is the failure mode the last section is about.

You replace each weight (originally a 16-bit float) with a small integer plus a per-group scale factor that says “multiply by this to get back to roughly the original number.” You then accept that the roughly part introduces a little noise, and you bet — correctly, most of the time — that the model has enough redundancy that the noise gets averaged out before it reaches the output.

How it works

Start with the dumbest version — round every weight in your 70B download to the nearest of 16 values — and watch what breaks. Each repair is one of the four ideas the whole field rests on.

Idea 1 — why the naive version doesn’t immediately explode: over-parameterization

A trained transformer has billions of weights, and individually most of them carry a tiny amount of information. The output of any given layer is a sum over thousands of weights times their inputs, and the intuition is that small rounding errors on each weight tend to average out before they reach the output — the central limit theorem is the shape of the argument, not a clean theorem about it (the errors aren’t IID and the model is learned, so the math isn’t strict). Round each weight to the nearest 4-bit value and most of the noise washes out inside that sum, most of the time.

The same fact shows up in the older neural-network literature as “overparameterization buys robustness” — pruning, distillation, and quantization all lean on it. Surveys like Gholami et al. (2021) frame the empirical version: networks tolerate weight noise far better than the bit-level math would predict.

A gap worth naming, and it is the field’s: there is no clean theory saying “a model with N parameters can survive log-X bits of weight noise per parameter.” The evidence is empirical — people quantize, measure the benchmark drop, and the published 4-bit results report small degradations relative to the 16-bit model. GPTQ’s paper is the standard citation, and the size of the drop varies enough by model, scheme and benchmark that a single number would mislead; read the tables instead. The intuition above is the right shape of explanation, but the precise envelope — including whether bigger models really tolerate more relative noise, or just more absolute parameters worth of noise — is something you measure, not derive.

Idea 2 — first repair: pick the range per group, not per model

The naive version has an obvious hole. Quantization works by picking a numerical range and slicing it into equal-sized buckets. Pick one range for the whole model and a single unusually large weight anywhere in it stretches that range, so every ordinary weight rounds to a coarser grid. If your weights were uniformly spread over a huge range, this scheme would be hopeless — most buckets would sit empty between extremes.

In trained transformers, weights mostly cluster tightly around zero in a roughly bell-shaped distribution. So a small fixed range (say, the 99.9th percentile of magnitudes in a layer) covers almost every weight, and you can give that range generous resolution. Grouped and per-channel scales — the group_size=128 setting you’ll see in GPTQ/AWQ implementations, which is an implementation convention rather than a claim from either paper — exploit this: cut the weights into small groups, pick a tight range per group, quantize inside that range. Each group gets its own scale factor, so a layer where one channel has wider weights doesn’t drag down the resolution of the rest. (GPTQ (Frantar et al., 2022) layers an additional trick on top — approximate second-order error compensation as it quantizes weight columns one at a time — but grouped scales are the part that the friendly distribution lets you get away with.)

This is the same reason JPEG works on photos: the signal is not adversarial, it has structure, and you can exploit the structure to spend bits where they matter. Where that analogy breaks: JPEG throws away detail your eye can’t resolve, and the loss is judged by a human looking at the result. Quantization has no equivalent perceptual model — nothing tells you in advance which weights the model can’t resolve, so which bits are safe to drop is an empirical question, settled by measuring, not by theory.

Idea 3 — why you’d bother: the bottleneck is bandwidth, not precision

Ideas 1 and 2 explain why quantization doesn’t hurt much. They don’t explain why it helps so much — after all, dequantizing costs extra work. That’s the part that makes quantization especially worth it for LLMs as opposed to, say, classical scientific computing.

When a dense transformer generates a token at small batch size, the GPU effectively has to read the model’s weights out of HBM into the compute units, multiply, and stream activations back. For a dense 70B model in FP16, that’s reading ~140 GB of weights per token. With H100’s ~3.35 TB/s of HBM bandwidth, the back-of-envelope roofline is ~40 ms per token even if the math itself took zero time — an order-of-magnitude sketch, not a benchmark. (Batching, mixture-of-experts sparsity, and KV-cache reads change the picture — but for batch-1 dense decode, the math is genuinely not the bottleneck.)

This is why “weight-only quantization” is the common trick in practice. You store the weights in INT4, you read them out of HBM at 4 bits per weight (4× less data), then you dequantize on the fly back to FP16 (or BF16) inside the kernel and do the actual matmul in the higher precision. You pay for some extra compute on dequantization — but compute was sitting idle anyway because you were waiting on memory. So the speedup tracks the bit reduction, up to what the kernel gives back in unpacking and non-weight traffic, and the quality cost is just the rounding error from idea 1.

This is the asymmetry that makes quantization shine for LLM inference specifically. In a compute-bound workload (like a small CNN doing image classification on a beefy GPU), shrinking the weights doesn’t help much because you weren’t bottlenecked on reading them. In LLM decode, you absolutely were, so the saving translates almost directly to throughput.

Idea 4 — the repair that per-group scales don’t cover: outliers

Per-group scales fix the weights. They don’t fix what the weights get multiplied by. If quantization were uniformly easy you’d see no papers about it; the interesting part is this failure mode, and it has a name: outlier features.

Tim Dettmers and collaborators showed in LLM.int8() (2022) that as transformers cross a certain scale (around 6.7B parameters in the models they studied), a small number of feature dimensions in the activations start carrying values much larger than the rest — the paper reports magnitudes up to ~20× the typical range, concentrated in specific dimensions. If you quantize naively, those outliers force you to pick a numerical range wide enough to contain them, which wastes precision on the overwhelming majority of “normal” values, and the model’s quality collapses.

LLM.int8() handles this by detecting those outlier dimensions at runtime and routing them through a 16-bit matrix multiplication while quantizing the rest to 8-bit. The paper reports keeping more than 99.9% of values in 8-bit while preserving accuracy on models up to 175B parameters. Later activation-aware schemes — SmoothQuant shifts the difficulty between weights and activations, AWQ chooses per-channel scales that protect the most salient channels — respond to the same underlying observation: not all weights and activations are created equal, and the few that matter most have to be protected.

This is the seam in the story. Quantization “just works” on the average weight; the engineering is almost entirely about the few percent of weights and activations where it doesn’t.

Where the seams show

A few honest caveats:

You started with quantization = low-precision weights + a calibrated map back. What did this post add? — + the reason the trade is lopsided for LLMs specifically. Rounding costs you a little accuracy because a trained network’s weights are redundant and clustered; it buys you a lot of speed because batch-1 decode is waiting on memory, not math. Change either half of that — a compute-bound workload, or a model whose weights aren’t redundant — and the same 4-bit download that made your 70B fit on one GPU stops being a good deal.

Check yourself

Before you go — you quantize two models to INT4: a 70B language model doing batch-1 chat, and a small image classifier running big batches on the same GPU. The language model gets dramatically faster. The classifier barely moves. Same bit reduction — why the different outcome?

Answer

Because they’re bottlenecked on different resources. Batch-1 decode has to stream every weight out of HBM to produce each token, so the runtime is set by bytes moved; cut the bytes 4× and you cut close to 4× of the wall clock. The classifier at large batch reuses each loaded weight across many inputs, so it’s limited by arithmetic throughput, not by loading — shrinking the weights doesn’t shrink the work that’s actually the bottleneck, and you’ve added dequantization work. The general rule: weight-only quantization pays off in proportion to how memory-bound you already were.

And one more: two teams quantize the same model to INT4. Team A uses one scale factor for each whole weight matrix; team B uses one per group of 128 weights. Team B’s model is noticeably better. Where did team A’s quality go, given that both used exactly 4 bits per weight?

Answer

Into the range. A single scale per matrix has to cover the largest magnitude anywhere in it, so one unusually large weight sets the spacing of the 16 buckets for every other weight in the matrix — and since weights cluster tightly near zero, most of them land in a couple of buckets near the middle and lose almost all their resolution. Group-wise scales let each group of 128 pick its own tight range, so the same 4 bits describe a much smaller interval. The bits didn’t change; what changed is how much numerical territory each bit had to cover. Note the cost: you now store a scale per group, so the true bits-per-weight is slightly above 4.

Going deeper