Why does GPU memory bandwidth matter more than FLOPS for LLM inference?
You bought the GPU for the teraflops. At inference time, almost none of them are doing anything. The bottleneck is moving the weights, not multiplying them.
On this page
The picture version
Six pictures for a reader who has never thought about what makes a chatbot fast. The prose below fills in the seams the pictures skip.
1 · The shape you already know
A stuttering video isn’t a slow computer. It’s a thin pipe.
2 · The wrong dial
You bought it for the arithmetic. That’s not what you’re using.
3 · What one word costs
Every weight rides past, gets used once, and is thrown away.
4 · The ceiling, in one division
Pipe speed divided by model size. That’s the whole estimate.
5 · The way out, and its limit
Run the belt once, serve thirty-two people.
6 · Keep this card
The whole thing on one index card.
Why it exists
Picture streaming a 4K movie on a slow Wi-Fi connection. Your laptop’s CPU is more than fast enough to play the video — it sits mostly idle. The movie stutters because frames can’t reach the laptop fast enough. Speeding up the CPU wouldn’t help; you need a fatter pipe. Running a large model on a GPU has exactly this shape — except that the GPU’s pipe isn’t the internet, it’s the few centimetres between the memory chips and the math units on the same board. The distance is trivial; the bandwidth is still the wall. The GPU has tens of thousands of math units doing nothing most of the time while the model’s weights — tens of gigabytes of them — are dragged across that gap fast enough to keep them fed. That’s memory bandwidth, and the running example for this post is a 70B model on a single H100.
If you’ve ever shopped for a GPU to run a local model, you’ve noticed something strange. Two cards can differ enormously in raw arithmetic throughput — their FLOPS rating — but when you’re generating from a 70B model for one user, the tokens per second track a different spec-sheet number much more closely: memory bandwidth, in GB/s.
A consumer RTX 4090 has roughly 1 TB/s of memory bandwidth; an H100 SXM has roughly 3.35 TB/s; Apple’s M2 Ultra has around 800 GB/s of unified memory. Those three numbers span a factor of about four. The compute numbers across the same three parts span a far wider range. If generating text were a compute problem, the ranking of these machines at single-user decoding should look like the compute spread — and it looks a lot more like the bandwidth spread instead.
The answer is the question this post exists to answer: while the model is writing out its answer one token at a time — the phase called decode — your GPU is barely computing anything. It is reading. For a dense model generating for a single user, producing one token means dragging essentially all of the model’s weights from VRAM to the compute units, doing a small amount of arithmetic, and discarding them; token N+1 reads them again. The throughput of that read pipe is the ceiling. (Batching and mixture-of-experts change this picture, and both show up later.)
Why it matters now
A surprising number of the cost and performance questions in LLM serving turn out to be memory-bandwidth questions wearing a disguise:
- “Why is my 7B model so much faster than my 70B model on the same GPU?” Roughly because there are 10× fewer bytes of weights to ship per token. The arithmetic is cheaper too, but the arithmetic was never the limit.
- “Why does quantization speed up inference so much, when it doesn’t reduce the number of multiplications?” Because cutting the bytes per weight cuts the bytes you have to move per token. Bandwidth-bound work scales with bytes, not with operations. (The saving is rarely the exact ratio — dequantization and packing cost something — but it’s the right direction and roughly the right size.)
- “Why does speculative decoding help so much?” Because verifying a short draft happens in one pass over the weights instead of one pass per token. You amortize the bandwidth bill over however many drafted tokens get accepted.
- “Why is batching the trick every serving engine reaches for?” Same reason. One weight read, many sequences using it.
- “Why are data-center GPUs so expensive when their FLOPS-per-dollar isn’t that extreme?” Partly because they ship HBM, which is expensive and supply-constrained in a way GDDR isn’t. It’s not the only cost driver, but it’s the one that maps directly onto the number this post is about.
If your mental cost model for inference is “FLOPS in, tokens out,” it is predicting the wrong things. A bandwidth-first model predicts the right things, and explains a pile of otherwise mysterious engineering choices.
The short answer
LLM decode speed ≈ GPU memory bandwidth ÷ model size in bytes
Picture to keep: a conveyor belt feeding a workshop of idle machinists. Every weight rides the belt past a machinist, gets used for exactly one multiply-and-add, and is thrown away. Tokens per second is how many times per second you can run the entire belt end to end.
To generate one token, the GPU has to stream every model weight from VRAM to its compute units. The compute itself is fast and finishes early; the weights take time to arrive. Tokens per second is, to a first approximation, how many times per second the GPU can drag the whole model across that pipe.
How it works
The intuitive cost model — “more FLOPS, more tokens per second” — is the thing to break first, and the fastest way is to price a single decode step (one new token) in each currency and see which one runs out.
Here’s what actually happens in that step:
- A small input — the new token’s vector — enters the model.
- For every layer, the GPU multiplies that vector by the layer’s weight matrices.
- The result becomes the input for the next layer.
- At the top, you sample a token.
Step 2 is “matrix times vector” — a huge grid of weights multiplied by one small vector, the current token’s activations. It has a property worth staring at: each weight is read from memory, used in one multiply-add, and then never touched again for this token.
That ratio — arithmetic per byte loaded — is called
arithmetic intensity.
Here it is about 1 operation per byte: one multiply-add (two operations, by
the usual convention) per 2-byte weight. Now compare against the chip. The
break-even point where a GPU stops being starved by memory is its peak
arithmetic rate divided by its bandwidth: for an H100 SXM that’s roughly
989 TFLOP/s ÷ 3.35 TB/s ≈ 300 operations per byte using its TF32
tensor-core figure, or about double that against its FP16 peak
(spec sheet). Decode
arrives with an intensity of about 1. It isn’t close, and it isn’t close by
two orders of magnitude. The compute units sit idle waiting for VRAM; the
bandwidth pipe runs flat out.
You can sanity-check this with a back-of-envelope calculation. A 70B-param model at FP16 is ~140 GB of weights. On a single H100 with ~3.35 TB/s of HBM bandwidth, the absolute ceiling on per-token decode is roughly:
3.35 TB/s ÷ 140 GB ≈ 24 tokens/sec
That’s a hard upper bound from physics, ignoring all overhead. Real systems land somewhere below it. (A 70B model doesn’t even fit on a single H100 in FP16, so in practice you’d shard or quantize, but the shape of the calculation is what matters.) Notice what’s not in that calculation: the GPU’s FLOPS rating. It doesn’t appear because it isn’t the bottleneck.
Now contrast with prefill — processing the prompt before any tokens come out. Prefill multiplies weights against many tokens at once (the whole prompt), so each weight gets reused across all those tokens. Arithmetic intensity goes up, the compute units actually get used, and prefill on a long prompt moves toward the opposite regime, where arithmetic rather than bandwidth is the limit. This is why “time to first token” (prefill-bound) and “tokens per second after that” (decode-bound) live on different curves and respond to different optimizations. The same GPU that is FLOPS-bound during prefill is bandwidth-bound during decode.
Why batching breaks the rule (and why it has limits)
Batching multiple users’ requests together is the closest thing to a free lunch the bandwidth model allows. If 32 users each want a token, one sweep of the weights can serve all 32 at once — the small vector becomes a small matrix, and each weight byte you moved does 32 multiply-adds instead of one. Arithmetic intensity goes up roughly 32×. Until, that is, something else fills up — usually the KV cache, which grows per request and per token and eventually pushes you back into a bandwidth wall, just on different data. Memory bandwidth is the dominant constraint; batching just lets you share the bill.
Where this gets fuzzy
- The ”≈” is doing real work. Real inference involves attention reads against the KV cache (also bandwidth-bound, but on cache bytes not weight bytes), kernel launch overhead, communication between GPUs in a sharded setup, and some genuinely compute-bound bits. The “bandwidth ÷ size” estimate is an upper bound, not a prediction.
- It’s a decode-time argument. Training, prefill, and very-large-batch serving all push toward compute-bound regimes. “Bandwidth matters more than FLOPS” is specifically a claim about single-request decode.
- Architectures change the picture. MoE models route each token through only some of the weights, so the bytes moved per token fall well below the total model size. How that plays out in a real deployment depends on routing, expert placement across GPUs, and batch composition — and the operators who run these at scale don’t generally publish those numbers, so treat any specific MoE bandwidth-versus-compute balance as unpublished rather than settled.
- Apple’s unified memory is a quieter but related story. One pool of memory, shared by CPU and GPU, with high bandwidth and large capacity — which is exactly the pair of properties single-user decoding wants. That explains why a Mac can hold and run a model whose size would need several discrete GPUs; it doesn’t make it competitive on the compute-heavy phases.
The point isn’t that FLOPS don’t matter. They matter for training, prefill, and image/video models with very different intensity profiles. The point is that the mental model “compute = speed” comes from a world that isn’t the one we’re in for LLM inference.
You started with LLM decode speed ≈ GPU memory bandwidth ÷ model size in bytes. What did the walk-through add? — + ...but only while arithmetic intensity stays near 1. That clause is the entire post: every trick in
production serving (batching, speculative decoding, MoE) is an attempt to
raise the number of useful tokens produced per trip down the belt, and every
one of them stops helping the moment something else — usually the KV cache —
becomes what you’re reading instead.
Check yourself
Before you go — you quantize a 70B model from FP16 to int4, cutting the weight bytes by 4×. Your single-user token rate improves a lot — not exactly 4×, but in that neighbourhood. Now you do the same thing on a server running a batch of 256 concurrent requests, and the speedup is much smaller. Why?
Answer
Because at batch 256 you’re no longer in the regime the compression line describes. Batching reuses one weight load across many sequences, so arithmetic intensity climbs and the workload drifts toward compute-bound — and int4 doesn’t cut the number of multiply-adds, only the bytes. Cheaper bytes stop buying speed once bytes stopped being the constraint. (Large batches also spend a growing share of their memory traffic on the KV cache, which quantizing the weights doesn’t touch.)
And: two cards have identical memory bandwidth, but one has twice the VRAM capacity. For serving one user a model that fits comfortably in both, does the bigger card generate tokens faster?
Answer
No — not for that one user. Decode speed is bandwidth ÷ bytes read per token, and neither term changed. What the extra capacity buys is concurrency: more requests’ KV caches resident at once, hence larger batches, hence more total tokens per second across users. Capacity raises throughput; bandwidth raises per-user latency. Confusing the two is why “it has 48 GB, it must be fast” disappoints people.
Famous related terms
- Arithmetic intensity —
arithmetic intensity = FLOPs ÷ bytes loaded— the dial that decides whether you’re compute-bound or memory-bound. - Roofline model —
roofline ≈ a chart with two ceilings: bandwidth and FLOPS— a one-page mental model from HPC for predicting which of the two is going to bite first. Worth knowing. - HBM —
HBM = stacked DRAM bonded to the GPU package— a major part of what makes a data-center GPU expensive, and the component whose supply the whole industry watches. - Quantization —
quantization = same weights, fewer bits each— speeds up decode because bytes-moved drops, not because operations drop. - KV cache — the other big bandwidth consumer at decode time; your GPU is reading both weights and cache on every token.
- MoE (Mixture of Experts) —
MoE = many experts, route each token to a few— changes the bytes-per-token math by activating only some weights. - Speculative decoding —
speculative decoding = small model drafts, big model verifies in parallel— works by amortizing one weight read across many tokens.
Going deeper
- Williams, Waterman & Patterson, Roofline: An Insightful Visual Performance Model (CACM, 2009) — the primary source for “is there a principled way to tell which ceiling I’m hitting?”, predating LLMs entirely; the decode-versus-prefill split above is an application of its framing, not a claim from the paper.
- Horace He, Making Deep Learning Go Brrrr From First Principles — read this if you want to know how to diagnose a real workload as memory-bound, compute-bound, or overhead-bound rather than just believe an argument about one.
- Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon et al., 2023) — the rabbit hole: what a serving system looks like once its designers have fully accepted that memory, not compute, is the scarce thing.