Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does GPU memory bandwidth matter more than FLOPS for LLM inference?

You bought the GPU for the teraflops. At inference time, almost none of them are doing anything. The bottleneck is moving the weights, not multiplying them.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who has never thought about what makes a chatbot fast. The prose below fills in the seams the pictures skip.

1 · The shape you already know

A stuttering video isn’t a slow computer. It’s a thin pipe.

4K movie, slow Wi-Fi the video out there thin your laptop: mostly idle A faster processor buys you nothing. the frames can’t arrive fast enough a big model on a GPU the weights in memory thin the math units: mostly idle A faster chip buys you nearly nothing. and this pipe is a few centimetres long Same shape. The distance is trivial; the pipe is still the wall.
Everyone has watched a video stutter on a machine that was doing nothing. The same thing happens inside a GPU running a large model — the arithmetic finishes early and waits for the next batch of numbers to arrive.

2 · The wrong dial

You bought it for the arithmetic. That’s not what you’re using.

arithmetic on the spec sheet a gaming card a data-centre card a desktop Mac an enormous spread memory bandwidth ~1 TB/s 3.35 TB/s 0.8 TB/s a gaming card a data-centre card a desktop Mac about a factor of four writing one user’s answer tracks this chart
If generating text were an arithmetic problem, these machines would rank the way the left chart ranks them. They rank much more like the right chart — which is the anomaly the rest of the post explains.

3 · What one word costs

Every weight rides past, gets used once, and is thrown away.

used once, then dropped tens of thousands of these, idle most of the time To write one word, the whole belt has to run end to end: 140 GB of weights, every time each weight does exactly one multiply-and-add before it’s discarded — so the arithmetic per byte moved is about as low as it can get the next word reads all 140 GB again
The huge grid of numbers is multiplied by one tiny vector, so nothing gets reused within a word. The machinists aren’t the bottleneck — the belt is, and the belt has to make a full pass per word.

4 · The ceiling, in one division

Pipe speed divided by model size. That’s the whole estimate.

3.35 TB/s ÷ 140 GB ≈ 24 words per second a hard ceiling from physics, before any real-world overhead — real systems land below it NOTICE WHAT ISN’T IN THIS SUM: THE CHIP’S ARITHMETIC RATING
Two numbers off the spec sheet predict the answer rate better than the number everyone shops for. The arithmetic rating doesn’t appear because it isn’t what runs out — it is idle by two orders of magnitude.

5 · The way out, and its limit

Run the belt once, serve thirty-two people.

one sweep of the weights every waiting user gets their word out of the same pass the bandwidth bill, split many ways until… the new thing being read each user’s conversation, growing every word Batching doesn’t repeal the rule. It changes what you spend the pipe on.
One weight load can serve many sequences, so the arithmetic per byte climbs and the arithmetic units finally get used. Then the stored conversations grow until they, not the weights, are what the pipe is carrying.

6 · Keep this card

The whole thing on one index card.

words per second ≈ how fast the pipe moves bytes ÷ how many bytes the model is …while each byte does about one operation every serving trick — batching, drafting ahead, routing to a slice of the model — is an attempt to raise that last number
Picture to keep: a conveyor belt feeding a workshop of idle machinists. Every weight rides past, gets used for exactly one multiply-and-add, and is thrown away; words per second is how often you can run the whole belt.

Why it exists

Picture streaming a 4K movie on a slow Wi-Fi connection. Your laptop’s CPU is more than fast enough to play the video — it sits mostly idle. The movie stutters because frames can’t reach the laptop fast enough. Speeding up the CPU wouldn’t help; you need a fatter pipe. Running a large model on a GPU has exactly this shape — except that the GPU’s pipe isn’t the internet, it’s the few centimetres between the memory chips and the math units on the same board. The distance is trivial; the bandwidth is still the wall. The GPU has tens of thousands of math units doing nothing most of the time while the model’s weights — tens of gigabytes of them — are dragged across that gap fast enough to keep them fed. That’s memory bandwidth, and the running example for this post is a 70B model on a single H100.

If you’ve ever shopped for a GPU to run a local model, you’ve noticed something strange. Two cards can differ enormously in raw arithmetic throughput — their FLOPS rating — but when you’re generating from a 70B model for one user, the tokens per second track a different spec-sheet number much more closely: memory bandwidth, in GB/s.

A consumer RTX 4090 has roughly 1 TB/s of memory bandwidth; an H100 SXM has roughly 3.35 TB/s; Apple’s M2 Ultra has around 800 GB/s of unified memory. Those three numbers span a factor of about four. The compute numbers across the same three parts span a far wider range. If generating text were a compute problem, the ranking of these machines at single-user decoding should look like the compute spread — and it looks a lot more like the bandwidth spread instead.

The answer is the question this post exists to answer: while the model is writing out its answer one token at a time — the phase called decode — your GPU is barely computing anything. It is reading. For a dense model generating for a single user, producing one token means dragging essentially all of the model’s weights from VRAM to the compute units, doing a small amount of arithmetic, and discarding them; token N+1 reads them again. The throughput of that read pipe is the ceiling. (Batching and mixture-of-experts change this picture, and both show up later.)

Why it matters now

A surprising number of the cost and performance questions in LLM serving turn out to be memory-bandwidth questions wearing a disguise:

If your mental cost model for inference is “FLOPS in, tokens out,” it is predicting the wrong things. A bandwidth-first model predicts the right things, and explains a pile of otherwise mysterious engineering choices.

The short answer

LLM decode speed ≈ GPU memory bandwidth ÷ model size in bytes

Picture to keep: a conveyor belt feeding a workshop of idle machinists. Every weight rides the belt past a machinist, gets used for exactly one multiply-and-add, and is thrown away. Tokens per second is how many times per second you can run the entire belt end to end.

To generate one token, the GPU has to stream every model weight from VRAM to its compute units. The compute itself is fast and finishes early; the weights take time to arrive. Tokens per second is, to a first approximation, how many times per second the GPU can drag the whole model across that pipe.

How it works

The intuitive cost model — “more FLOPS, more tokens per second” — is the thing to break first, and the fastest way is to price a single decode step (one new token) in each currency and see which one runs out.

Here’s what actually happens in that step:

  1. A small input — the new token’s vector — enters the model.
  2. For every layer, the GPU multiplies that vector by the layer’s weight matrices.
  3. The result becomes the input for the next layer.
  4. At the top, you sample a token.

Step 2 is “matrix times vector” — a huge grid of weights multiplied by one small vector, the current token’s activations. It has a property worth staring at: each weight is read from memory, used in one multiply-add, and then never touched again for this token.

That ratio — arithmetic per byte loaded — is called arithmetic intensity. Here it is about 1 operation per byte: one multiply-add (two operations, by the usual convention) per 2-byte weight. Now compare against the chip. The break-even point where a GPU stops being starved by memory is its peak arithmetic rate divided by its bandwidth: for an H100 SXM that’s roughly 989 TFLOP/s ÷ 3.35 TB/s ≈ 300 operations per byte using its TF32 tensor-core figure, or about double that against its FP16 peak (spec sheet). Decode arrives with an intensity of about 1. It isn’t close, and it isn’t close by two orders of magnitude. The compute units sit idle waiting for VRAM; the bandwidth pipe runs flat out.

You can sanity-check this with a back-of-envelope calculation. A 70B-param model at FP16 is ~140 GB of weights. On a single H100 with ~3.35 TB/s of HBM bandwidth, the absolute ceiling on per-token decode is roughly:

3.35 TB/s ÷ 140 GB ≈ 24 tokens/sec

That’s a hard upper bound from physics, ignoring all overhead. Real systems land somewhere below it. (A 70B model doesn’t even fit on a single H100 in FP16, so in practice you’d shard or quantize, but the shape of the calculation is what matters.) Notice what’s not in that calculation: the GPU’s FLOPS rating. It doesn’t appear because it isn’t the bottleneck.

Now contrast with prefill — processing the prompt before any tokens come out. Prefill multiplies weights against many tokens at once (the whole prompt), so each weight gets reused across all those tokens. Arithmetic intensity goes up, the compute units actually get used, and prefill on a long prompt moves toward the opposite regime, where arithmetic rather than bandwidth is the limit. This is why “time to first token” (prefill-bound) and “tokens per second after that” (decode-bound) live on different curves and respond to different optimizations. The same GPU that is FLOPS-bound during prefill is bandwidth-bound during decode.

Why batching breaks the rule (and why it has limits)

Batching multiple users’ requests together is the closest thing to a free lunch the bandwidth model allows. If 32 users each want a token, one sweep of the weights can serve all 32 at once — the small vector becomes a small matrix, and each weight byte you moved does 32 multiply-adds instead of one. Arithmetic intensity goes up roughly 32×. Until, that is, something else fills up — usually the KV cache, which grows per request and per token and eventually pushes you back into a bandwidth wall, just on different data. Memory bandwidth is the dominant constraint; batching just lets you share the bill.

Where this gets fuzzy

The point isn’t that FLOPS don’t matter. They matter for training, prefill, and image/video models with very different intensity profiles. The point is that the mental model “compute = speed” comes from a world that isn’t the one we’re in for LLM inference.

You started with LLM decode speed ≈ GPU memory bandwidth ÷ model size in bytes. What did the walk-through add? — + ...but only while arithmetic intensity stays near 1. That clause is the entire post: every trick in production serving (batching, speculative decoding, MoE) is an attempt to raise the number of useful tokens produced per trip down the belt, and every one of them stops helping the moment something else — usually the KV cache — becomes what you’re reading instead.

Check yourself

Before you go — you quantize a 70B model from FP16 to int4, cutting the weight bytes by 4×. Your single-user token rate improves a lot — not exactly 4×, but in that neighbourhood. Now you do the same thing on a server running a batch of 256 concurrent requests, and the speedup is much smaller. Why?

Answer

Because at batch 256 you’re no longer in the regime the compression line describes. Batching reuses one weight load across many sequences, so arithmetic intensity climbs and the workload drifts toward compute-bound — and int4 doesn’t cut the number of multiply-adds, only the bytes. Cheaper bytes stop buying speed once bytes stopped being the constraint. (Large batches also spend a growing share of their memory traffic on the KV cache, which quantizing the weights doesn’t touch.)

And: two cards have identical memory bandwidth, but one has twice the VRAM capacity. For serving one user a model that fits comfortably in both, does the bigger card generate tokens faster?

Answer

No — not for that one user. Decode speed is bandwidth ÷ bytes read per token, and neither term changed. What the extra capacity buys is concurrency: more requests’ KV caches resident at once, hence larger batches, hence more total tokens per second across users. Capacity raises throughput; bandwidth raises per-user latency. Confusing the two is why “it has 48 GB, it must be fast” disappoints people.

Going deeper