Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why is the KV cache a thing?

The model has to read your whole prompt every time it picks a token. Why doesn't it choke? Because of a quiet trick almost nobody mentions in the docs.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who has never wondered how a chat reply streams. The prose below fills in the seams the pictures skip.

1 · The problem

Every word of the answer means re-reading the whole conversation.

your 2,000-token conversation your 2,000-token conversation your 2,000-token conversation word 1 word 2 word 3 … 500 times, for a 500-word answer and the work in one pass grows with the square of the length this should crawl. it doesn’t.
To write the second word the model looks at your whole conversation plus the first word; for the third, plus the first two. Taken literally, a long chat should stall — instead the words arrive faster than you read them.

2 · The waste

Almost all of that work is the same work, again.

for each token, the layer works out two things: what it advertises and what it contributes step 1 step 2 step 3 recomputed, identical every time the only new one SAME NUMBERS, 499 TIMES OVER Each step adds one genuinely new column and redoes every old one.
The values a given word contributes to attention depend only on that word and what came before it. Recomputing them at every step is pure repetition — the same arithmetic producing the same numbers.

3 · Why that’s safe to reuse

Words can only look backwards. The past never gets edited.

word 1 word 2 word 3 word 4 the new word it reads all of them no arrow runs this way — nothing later can change what an earlier word offers
Attention in these models is one-directional: each position can read earlier positions and never later ones. So once a word’s contribution has been worked out, it is fixed for the rest of the answer — which is exactly the licence to keep it instead of redoing it.

4 · The fix

Everyone already seated keeps holding up their card.

name + note name + note name + note name + note already computed · never rewritten · just held up the new word reads every card on the table — one cheap sweep writes its own card Per word: one word’s worth of new work + a read over the cards. the first pass over your prompt (“prefill”) still does the full expensive read — that’s the pause before the answer starts. everything after it is the cheap phase.
Keep every earlier word’s two values in memory and the per-word cost collapses from “redo everything” to “do one word, then glance at the table.” This is why time-to-first-word and words-per-second afterwards are reported as separate numbers.

5 · The bill

You didn’t delete the work. You turned it into a memory bill.

140 GB the model itself 42 GB one conversation’s cards, at the longest context it allows 160 GB — all the memory two GPUs have 182 GB together doesn’t fit … and that’s one user. a server holds one set of cards per chat. figures for a 70-billion-parameter open model — the arithmetic is in the prose
Roughly a third of a megabyte per word, per conversation, held for as long as the chat lives. That is why so much serving engineering — batching, paging, cache quantization, shared key/value heads — exists to squeeze this one number.

6 · Keep this card

The whole trick on one index card.

KV cache = keep what every earlier word offers + compute only the new word + pay for it in memory, not time — and it only works because words look backwards
Picture to keep: a long conference table where everyone seated holds up a name card and a note, and only the new arrival has to write one. The cards aren’t free to hold up — they take up the table, and the table is what runs out first.

Why it exists

You’re forty messages deep in a chat. You paste in a long document and ask a question, there’s a pause, and then the answer streams out faster than you can read it. Notice how strange that is: the conversation above your question is thousands of tokens long, and the model is supposedly re-reading all of it to produce each word.

That’s the running example for this post — a 2,000-token prompt producing a 500-token answer — and here’s why it should bother you. An LLM takes your prompt and produces one next token. To produce the second token it looks at the prompt plus that first generated token. To produce the third, the prompt plus the first two. Every step extends the input by one token and runs the whole forward pass again.

Taken literally, your 500-token answer means 500 forward passes over sequences of length 2,001, 2,002, 2,003, … 2,500 — and each pass is, naively, quadratic in sequence length because of attention. That should be unusably slow. Instead tokens arrive at tens per second. How?

The answer is the KV cache. Every open inference stack you can read is built around it, and it rarely shows up in introductions. It’s why your tokens stream instead of stalling. It’s also why “context window” has a memory cost, why long prompts get expensive in a non-obvious way, and why the GPU serving you usually runs out of its own memory before it runs out of arithmetic.

Why it matters now

KV cache is load-bearing in a way most people don’t notice until something breaks:

If you’re building anything that calls an LLM at scale, the cost model in your head should have “KV cache” in it.

The short answer

KV cache = a per-layer store of the key and value vectors for every token already in the sequence, reused so attention only has to compute K and V for the new token

Picture to keep: a long conference table where each seated person holds up a name card and a note. A new arrival reads every card and note before speaking — but nobody already seated ever rewrites theirs. Except that the cards aren’t free to hold up: they occupy the table, and the table is what runs out first.

A transformer’s attention layer computes three vectors per token: Q, K, and V. Q is what this token is asking about; K and V are what every token offers up to be attended to. The crucial asymmetry: when you generate token N+1, the K and V for tokens 1…N can’t change, because attention only looks backwards. They were already computed. The cache just keeps them around so you don’t redo the work.

How it works

Start from the naive loop and let each cost force the next move. Here’s transformer inference written out honestly:

prompt -> forward pass -> next token
prompt + token1 -> forward pass -> token2
prompt + token1 + token2 -> forward pass -> token3
...

Inside each forward pass, every attention layer does, for each input token, something like:

Q_i = x_i · W_Q
K_i = x_i · W_K
V_i = x_i · W_V
attention_i = softmax(Q_i · K_all / sqrt(d)) · V_all

Why the naive loop breaks: at step 500 you recompute K_i and V_i for all 2,499 earlier tokens, having computed exactly the same numbers 499 times already. Two things make that redundancy visible:

  1. K_i and V_i depend on the hidden state at position i — which already folds in every earlier token — but they don’t depend on anything that comes after, because attention is causally masked. Appending token 2,501 cannot change them. So once computed for a given position, they’re frozen for the rest of generation.
  2. To compute attention for the new token, you need the full K_all and V_all — the keys and values for every position so far. But you already computed almost all of them on previous steps.

Fix: keep K and V for every token, every layer, in GPU memory. On each new generation step:

That changes the per-step cost from “redo everything” to “do one token’s worth of compute, plus a single attention read against the cache.” The quadratic blow-up is gone for the generation phase.

There are two distinct phases people sometimes conflate:

That split is why “time to first token” and “tokens per second after the first” are reported separately. Prefill cost lives in the first; cache amortization lives in the second.

What it costs

But you didn’t remove the work — you moved it into memory. That’s the trade the whole rest of modern inference is organized around. The cache size, per request, is roughly:

2 (K and V) × num_layers × num_kv_heads × head_dim × seq_len × bytes_per_element

Put your 2,500-token conversation through it with Llama 3.1 70B’s published config (80 layers, 8 KV heads, head dim 128, bf16): that’s 2 × 80 × 8 × 128 × 2 bytes ≈ 0.33 MB per token, so ~0.8 GB for this one chat. Push the same request to a 128k-token context and it’s ~42 GB. For scale: the weights alone are ~140 GB in bf16, so this model is already spread across at least two 80 GB H100 GPUs — and 140 GB of weights plus one maxed-out request’s 42 GB does not fit in that pair’s 160 GB. Now multiply the cache by however many requests you’re serving concurrently. This is why batching, paging (PagedAttention), quantization of the cache, and architectures like grouped-query attention exist. A large share of the engineering in modern inference engines is squeezing this one number.

Where it gets subtle

You started with KV cache = a per-layer store of K and V for every token already in the sequence. What did the walk-through add? — + it only works because attention is causal, and + you didn't delete the work, you converted it into a memory bill. The first is the precondition nobody states; the second is why, for most serving setups, it’s GPU memory rather than raw arithmetic that decides how many conversations you can hold at once.

Check yourself

Before you go — two models are the same size, same layer count, same head dimension. One uses grouped-query attention with 8 KV heads; the other uses full multi-head attention with 64. Serving the same long conversations on one GPU, how does the number of concurrent users you can hold differ, and why?

Answer

The GQA model holds roughly eight times as many. num_kv_heads sits directly in the cache-size formula, so 64 KV heads means 8× the cache bytes per token per request. Weights are about the same and per-token compute barely differs — but each concurrent request now reserves 8× the VRAM for its cache, so you run out of table long before you run out of math. That pressure is exactly why grouped-query attention exists.

And: you’re serving an API where every request begins with the same 8,000-token system prompt, then a short user question. Which phase does prompt caching help — prefill or decode — and roughly how much of the request does it eliminate?

Answer

Prefill, and most of it. Prefill is the expensive one-time pass over the prompt; decode is already cheap per step thanks to the cache. If the first 8,000 tokens are byte-identical across requests, their K and V are identical too, so the server can keep that prefix’s cache resident and start prefill at token 8,001. The saving is a prefill saving — you’d see it in time-to-first-token, not in tokens-per-second afterwards. And it evaporates the moment anything earlier in the prefix changes, because everything after an edit has to be recomputed.

Going deeper