Why quantization works
Stuffing a 70-billion-parameter model into 4-bit weights sounds like it should ruin it. It mostly doesn't — and the reason is more about how the model gets used at inference than about the math of rounding.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Idea 1 — why the naive version doesn’t immediately explode: over-parameterization
- Idea 2 — first repair: pick the range per group, not per model
- Idea 3 — why you’d bother: the bottleneck is bandwidth, not precision
- Idea 4 — the repair that per-group scales don’t cover: outliers
- Where the seams show
- Check yourself
- Famous related terms
- Going deeper
The picture version
Six pictures for a reader who has only ever seen the file-size column. The prose below fills in the seams the pictures skip.
1 · The problem
Cut every number to a quarter of its precision. Why isn’t it ruined?
2 · Why it survives at all
Each error is tiny and there are billions of them.
3 · First repair
Pick the range per small group, not once for the whole model.
4 · Why you’d bother
The GPU was waiting on memory, not on maths.
5 · Where it actually breaks
A handful of dimensions refuse to be average.
6 · Keep this card
The whole thing on one index card.
Why it exists
You go to download an open-weights model to run on your own machine. The
full 70-billion-parameter release is around 140 GB of raw weights and
your GPU has 24. Then you notice the file list has half a dozen other
versions of the same model — Q4, Q5, Q8 — and the 4-bit one is
roughly a quarter the size (a bit more, once you count the metadata each
scheme carries), and the comments underneath say it’s basically fine. That’s a 4× cut in the
precision of every single number in the model. If you did that to a
photograph, it would visibly fall apart. If you did it to audio, it would
hiss. Why doesn’t the model just become gibberish?
It’s a fair question, and the answer is the whole reason those Q4 files
exist at all. The short version: a trained
LLM
turns out to be much more robust to weight noise than its precision
suggests, and inference turns out to be bottlenecked on something
quantization directly relieves — moving bytes, not crunching numbers.
So the cost of quantization is small and the benefit is large, in a
way that’s specific to how this kind of model gets used, not a
universal law about neural networks.
This post is about why that asymmetry exists. The mechanics of any particular scheme (GPTQ, AWQ, bitsandbytes, FP8, INT4, MXFP4) you can look up; the load-bearing intuition is what makes them all keep working.
Why it matters now
Quantization is the lever that decides whether a given model actually runs on a given GPU. A 70B model in 16-bit weights is roughly 140 GB. An H100 SXM has 80 GB of VRAM (post). The model does not fit. In INT8 the weights are around 70 GB — technically under the line, but weights aren’t the only thing that has to live there, so in practice you have almost nothing left for KV cache or activations. In INT4 they’re around 35 GB and you have real headroom for users. The whole “can a hobbyist run this on one card?” question is decided here.
The same lever shows up in serving economics. Dense decode at small batch is memory-bandwidth bound — the GPU has to stream the model’s weights through the compute units to produce each token — so cutting weight bytes by 4× cuts the per-token weight traffic by up to 4×. Treat that as a ceiling, not a promise: scale metadata, non-weight traffic, and kernel quality all eat into it. Still, it’s why quantized builds are the norm rather than the exception in on-device runtimes (llama.cpp, MLX) and why serving stacks ship quantization support out of the box. NVIDIA’s Hopper generation even introduced FP8 in hardware so that training could move below 16-bit (NVIDIA H100 Transformer Engine announcement).
The short answer
quantization = lower-precision storage of weights + a calibrated map back to real numbers + a quality budget you spend carefully
Picture to keep: a wall of dimmer switches. Instead of a knob that slides anywhere, each one now clicks into 16 detents — but the electrician sets the range of each panel of dimmers to match how bright those lamps actually get, so almost nobody ends up at the wrong brightness. Where the analogy breaks: a room with one wrong dimmer just looks slightly off, whereas a handful of specific weights and activations matter far more than the rest — which is the failure mode the last section is about.
You replace each weight (originally a 16-bit float) with a small integer plus a per-group scale factor that says “multiply by this to get back to roughly the original number.” You then accept that the roughly part introduces a little noise, and you bet — correctly, most of the time — that the model has enough redundancy that the noise gets averaged out before it reaches the output.
How it works
Start with the dumbest version — round every weight in your 70B download to the nearest of 16 values — and watch what breaks. Each repair is one of the four ideas the whole field rests on.
Idea 1 — why the naive version doesn’t immediately explode: over-parameterization
A trained transformer has billions of weights, and individually most of them carry a tiny amount of information. The output of any given layer is a sum over thousands of weights times their inputs, and the intuition is that small rounding errors on each weight tend to average out before they reach the output — the central limit theorem is the shape of the argument, not a clean theorem about it (the errors aren’t IID and the model is learned, so the math isn’t strict). Round each weight to the nearest 4-bit value and most of the noise washes out inside that sum, most of the time.
The same fact shows up in the older neural-network literature as “overparameterization buys robustness” — pruning, distillation, and quantization all lean on it. Surveys like Gholami et al. (2021) frame the empirical version: networks tolerate weight noise far better than the bit-level math would predict.
A gap worth naming, and it is the field’s: there is no clean theory saying “a model with N parameters can survive log-X bits of weight noise per parameter.” The evidence is empirical — people quantize, measure the benchmark drop, and the published 4-bit results report small degradations relative to the 16-bit model. GPTQ’s paper is the standard citation, and the size of the drop varies enough by model, scheme and benchmark that a single number would mislead; read the tables instead. The intuition above is the right shape of explanation, but the precise envelope — including whether bigger models really tolerate more relative noise, or just more absolute parameters worth of noise — is something you measure, not derive.
Idea 2 — first repair: pick the range per group, not per model
The naive version has an obvious hole. Quantization works by picking a numerical range and slicing it into equal-sized buckets. Pick one range for the whole model and a single unusually large weight anywhere in it stretches that range, so every ordinary weight rounds to a coarser grid. If your weights were uniformly spread over a huge range, this scheme would be hopeless — most buckets would sit empty between extremes.
In trained transformers, weights mostly cluster tightly around zero
in a roughly bell-shaped distribution. So a small fixed range
(say, the 99.9th percentile of magnitudes in a layer) covers almost
every weight, and you can give that range generous resolution. Grouped
and per-channel scales — the group_size=128 setting you’ll see in
GPTQ/AWQ implementations, which is an implementation convention rather
than a claim from either paper — exploit this: cut the weights into small
groups, pick a tight range per group, quantize inside that range. Each group
gets its own scale factor, so a layer where one channel has wider
weights doesn’t drag down the resolution of the rest. (GPTQ
(Frantar et al., 2022) layers an
additional trick on top — approximate second-order error
compensation as it quantizes weight columns one at a time — but
grouped scales are the part that the friendly distribution lets you
get away with.)
This is the same reason JPEG works on photos: the signal is not adversarial, it has structure, and you can exploit the structure to spend bits where they matter. Where that analogy breaks: JPEG throws away detail your eye can’t resolve, and the loss is judged by a human looking at the result. Quantization has no equivalent perceptual model — nothing tells you in advance which weights the model can’t resolve, so which bits are safe to drop is an empirical question, settled by measuring, not by theory.
Idea 3 — why you’d bother: the bottleneck is bandwidth, not precision
Ideas 1 and 2 explain why quantization doesn’t hurt much. They don’t explain why it helps so much — after all, dequantizing costs extra work. That’s the part that makes quantization especially worth it for LLMs as opposed to, say, classical scientific computing.
When a dense transformer generates a token at small batch size, the GPU effectively has to read the model’s weights out of HBM into the compute units, multiply, and stream activations back. For a dense 70B model in FP16, that’s reading ~140 GB of weights per token. With H100’s ~3.35 TB/s of HBM bandwidth, the back-of-envelope roofline is ~40 ms per token even if the math itself took zero time — an order-of-magnitude sketch, not a benchmark. (Batching, mixture-of-experts sparsity, and KV-cache reads change the picture — but for batch-1 dense decode, the math is genuinely not the bottleneck.)
This is why “weight-only quantization” is the common trick in practice. You store the weights in INT4, you read them out of HBM at 4 bits per weight (4× less data), then you dequantize on the fly back to FP16 (or BF16) inside the kernel and do the actual matmul in the higher precision. You pay for some extra compute on dequantization — but compute was sitting idle anyway because you were waiting on memory. So the speedup tracks the bit reduction, up to what the kernel gives back in unpacking and non-weight traffic, and the quality cost is just the rounding error from idea 1.
This is the asymmetry that makes quantization shine for LLM inference specifically. In a compute-bound workload (like a small CNN doing image classification on a beefy GPU), shrinking the weights doesn’t help much because you weren’t bottlenecked on reading them. In LLM decode, you absolutely were, so the saving translates almost directly to throughput.
Idea 4 — the repair that per-group scales don’t cover: outliers
Per-group scales fix the weights. They don’t fix what the weights get multiplied by. If quantization were uniformly easy you’d see no papers about it; the interesting part is this failure mode, and it has a name: outlier features.
Tim Dettmers and collaborators showed in LLM.int8() (2022) that as transformers cross a certain scale (around 6.7B parameters in the models they studied), a small number of feature dimensions in the activations start carrying values much larger than the rest — the paper reports magnitudes up to ~20× the typical range, concentrated in specific dimensions. If you quantize naively, those outliers force you to pick a numerical range wide enough to contain them, which wastes precision on the overwhelming majority of “normal” values, and the model’s quality collapses.
LLM.int8() handles this by detecting those outlier dimensions at runtime and routing them through a 16-bit matrix multiplication while quantizing the rest to 8-bit. The paper reports keeping more than 99.9% of values in 8-bit while preserving accuracy on models up to 175B parameters. Later activation-aware schemes — SmoothQuant shifts the difficulty between weights and activations, AWQ chooses per-channel scales that protect the most salient channels — respond to the same underlying observation: not all weights and activations are created equal, and the few that matter most have to be protected.
This is the seam in the story. Quantization “just works” on the average weight; the engineering is almost entirely about the few percent of weights and activations where it doesn’t.
Where the seams show
A few honest caveats:
- Quantization is post-hoc, mostly. The recipe this post describes is: train in BF16 or FP16, then quantize for inference. There are quantization-aware training schemes too, and FP8 training is real on Hopper hardware. “Train high, serve low” still looks like the common shape, but the labs don’t publish their production precision recipes, so there is no number to put on it. Either way, the story above is about that asymmetry; training has its own constraints.
- Aggressive quantization eventually does break. 4-bit with careful schemes is where the published results cluster; below that, quality loss shows up faster, and 2-bit and 1-bit weights remain a research area rather than a default. The cliff exists; you just hit it lower than your intuition suggested, and exactly where depends on the model, the scheme, and the benchmark.
- Benchmarks can mislead about quality. A model can score similarly on MMLU after quantization but feel worse on long-context tasks, multilingual tasks, or code generation. The quality budget isn’t a single number — which means “we quantized and the benchmark held” is a weaker statement than it sounds, and reading it as “no users will notice” is a bet, not a measurement.
- The KV cache is also being quantized. Most of this post is about weight quantization, but serving engines also offer a lower-precision KV cache (FP8 is the common option) to fit longer contexts in the same memory. The mechanisms are similar but the failure modes differ — KV outliers behave differently from weight outliers.
You started with quantization = low-precision weights + a calibrated map back. What did this post add? — + the reason the trade is lopsided for LLMs specifically. Rounding costs you a little accuracy because a
trained network’s weights are redundant and clustered; it buys you a lot
of speed because batch-1 decode is waiting on memory, not math. Change
either half of that — a compute-bound workload, or a model whose weights
aren’t redundant — and the same 4-bit download that made your 70B fit on
one GPU stops being a good deal.
Check yourself
Before you go — you quantize two models to INT4: a 70B language model doing batch-1 chat, and a small image classifier running big batches on the same GPU. The language model gets dramatically faster. The classifier barely moves. Same bit reduction — why the different outcome?
Answer
Because they’re bottlenecked on different resources. Batch-1 decode has to stream every weight out of HBM to produce each token, so the runtime is set by bytes moved; cut the bytes 4× and you cut close to 4× of the wall clock. The classifier at large batch reuses each loaded weight across many inputs, so it’s limited by arithmetic throughput, not by loading — shrinking the weights doesn’t shrink the work that’s actually the bottleneck, and you’ve added dequantization work. The general rule: weight-only quantization pays off in proportion to how memory-bound you already were.
And one more: two teams quantize the same model to INT4. Team A uses one scale factor for each whole weight matrix; team B uses one per group of 128 weights. Team B’s model is noticeably better. Where did team A’s quality go, given that both used exactly 4 bits per weight?
Answer
Into the range. A single scale per matrix has to cover the largest magnitude anywhere in it, so one unusually large weight sets the spacing of the 16 buckets for every other weight in the matrix — and since weights cluster tightly near zero, most of them land in a couple of buckets near the middle and lose almost all their resolution. Group-wise scales let each group of 128 pick its own tight range, so the same 4 bits describe a much smaller interval. The bits didn’t change; what changed is how much numerical territory each bit had to cover. Note the cost: you now store a scale per group, so the true bits-per-weight is slightly above 4.
Famous related terms
- GPTQ —
GPTQ = post-training INT4 weight quantization + second-order error correction per layer— the workhorse 4-bit method on open-weight models; from Frantar et al., 2022. - AWQ —
AWQ = activation-aware weight quantization + per-channel scaling that protects salient weights— alternative to GPTQ that uses activation statistics rather than Hessian information. - LLM.int8() —
LLM.int8() = 8-bit matmul + a 16-bit side-channel for outlier feature dimensions— the paper that named the outlier problem and made INT8 inference work at 175B scale. - FP8 (E4M3 / E5M2) —
FP8 = 8-bit floating point + two formats for activations vs gradients— Hopper’s hardware bet that even training can move below 16 bits; see NVIDIA Transformer Engine docs. - BF16 vs FP16 —
BF16 = FP32's exponent + truncated mantissa— the precision most modern training is already done at; quantization compresses below this. - Weight-only quantization —
weight-only quantization = compressed weights in HBM + dequantize-on-the-fly inside the kernel— the variant that exploits the memory-bandwidth bottleneck specifically.
Going deeper
- Dettmers, Lewis, Belkada, Zettlemoyer — LLM.int8(): 8-bit Matrix Multiplication for Transformers at Scale (NeurIPS 2022). arXiv. Read this for the evidence behind idea 4 — what outlier features actually look like, at what scale they appear, and why they break naive quantization.
- Gholami et al. — A Survey of Quantization Methods for Efficient Neural Network Inference (2021). arXiv. The explainer: a readable tour of the whole landscape, and the best single answer to “where does the claim that networks tolerate weight noise actually come from?”
- Frantar, Ashkboos, Hoefler, Alistarh — GPTQ (ICLR 2023). arXiv. The rabbit hole: how much extra quality you can recover by choosing rounding errors deliberately instead of rounding each weight independently.