Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does prompt caching exist?

Your agent sends the same 50,000-token system prompt on every turn. Providers charge a fraction of the usual rate when they recognize it — not out of generosity, but because they stopped doing the work.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has never looked at an inference bill. The prose below fills in the seams the pictures skip.

1 · The problem

99% of what you send is what you already sent.

turn 1 system prompt + tool definitions + history turn 2 token-for-token identical to turn 1 turn 3 still identical the only new part And the server re-runs the full 50,000-token prefill each time. three seconds after doing exactly that work and throwing the result away
Prefill — the first pass over the whole prompt — is the most expensive part of inference on a long prompt. Both sides know the work is duplicated, which is why the fix is less an optimisation than an obvious correction.

2 · The wrong cache

Caching the answer saves nothing here.

the cache you’d reach for same input → remember the output useless in an agent loop why it can’t work the next request isn’t the same input — it’s the same 50,000 tokens plus a new one and you want fresh generation anyway So cache the intermediate state, not the answer. the KV cache for tokens 1…N is exactly the artifact prefill produced — keep it and the next request can attend against it directly prefill on N tokens is roughly N² in the attention layers; reusing it, you only pay for the new M-token suffix
When M is 200 new tokens against a 50,000-token prefix, that is an enormous saving — and the quadratic explosion in N disappears entirely. Prefill is what gets reused; decode is what you still pay for fresh.

3 · How the server recognises you

It hashes the prefix in blocks. One token differs, everything after diverges.

last turn’s prefix, hashed in blocks all cached this turn, with a timestamp added at the top all missed one token changed here … and every block after it hashes differently … “Exact prefix match” is not marketing simplification. It is load-bearing. a token’s cached keys and values were computed by attending over everything before it — change token 5 and every entry after describes a different context block hashing is how the open implementations do it; the closed providers don’t publish their internals
Caches are also scoped per organisation, so two customers who happen to share a prefix don’t share a cache. Insert a timestamp, reorder your tool list, or change one token of whitespace and you pay full price.

4 · The foot-gun

Miss every time and the bill goes up.

cache read ~10% of base base input what you’d pay uncached cache write 125% of base (5 min) 200% (1 hour) Mark a breakpoint on something that changes and you only ever write.
Anthropic’s ratios, as an example of the shape — the multipliers differ by provider and drift over time, so treat the numbers as a snapshot. The durable part is the ordering: read < base < write, with a write only worth it if you will hit it more than about once.

5 · Why it expires so fast

Five minutes is an eviction policy wearing a feature’s clothes.

one user’s 50,000-token prefix gigabytes, not megabytes of the most expensive memory in the building now multiply by every user keeping them all resident tiles out the GPUs immediately So the TTL is short because the memory is precious, not because five is a nice number. DeepSeek bets differently and backs its cache with disk: cheaper to retain, slower to fetch disk wins for long-tail prefixes hit hours later; VRAM wins for chatty agent loops where the next request lands in two seconds
Exact per-request cache size depends on the model’s layer count, head dimensions and cache precision, and providers don’t publish theirs — but the order of magnitude is what drives the design. A longer TTL costs more because the memory it pins is the scarce thing.

6 · Keep this card

The whole thing on one index card.

prompt caching = the KV cache for the prefix + persisted across requests + matched by exact-prefix hashing ∴ one timestamp at the top flips discount into premium
Picture to keep: the model reading your 50,000-token prefix is a clerk pulling the same fat case file off the shelf and re-reading it cover to cover for every question. Caching is leaving the file open on the desk between questions — and it gets cleared off if nobody comes back within a few minutes. Where it breaks: a clerk would recognise a file with one line edited. This matching is on the exact token prefix.

Why it exists

Build an agent for five minutes and you’ll notice something uncomfortable.

Every turn, you’re sending the model the same long system prompt, the same tool definitions, the same scrollback of “here’s what we did last step.” The variable part — the user’s new message, or the result of the last tool call — is tiny. Maybe 1% of the input tokens. The other 99% is identical to what you sent on the previous request.

The provider, on the other hand, is in an even more uncomfortable position. From their side: a request comes in with 50,000 tokens of prefix. Their server runs the full prefill pass — quadratic in prompt length, the most expensive thing the model does — to populate the KV cache for those 50,000 tokens. The model emits a few hundred tokens of output. The request ends. The KV cache is freed. Three seconds later, the same client sends 50,001 tokens — the original 50,000 plus one new user message. The server does the entire 50,000-token prefill again, from scratch, because nothing was kept around.

That is an absurd amount of duplicated work. Both sides know it. Prompt caching is the obvious move: keep the KV cache for the prefix around between requests (in GPU memory, or on a slower tier behind it), and bill the user for the cache hit instead of the recomputation. Prefill is the most expensive part of inference for long prompts. Skipping it is what pays for the discount on the pricing page — on Anthropic’s API, a cache hit bills at roughly 10% of the base input rate. (Discounts differ by provider; the shape doesn’t.)

Why it matters now

A few years ago, prompts were short. “Summarize this paragraph.” Caching across requests would have saved nothing worth the engineering. Three things changed:

In all three cases, the prefix is stable and the suffix is small. That’s the exact regime where caching wins. Without it, agentic workloads would be priced as if every step is a fresh request — which is to say, mostly priced as prefill cost on a prompt that was already prefilled five seconds ago.

It also matters because the cost model is not symmetric. On Anthropic’s API, a cache hit is priced at roughly 10% of the base input rate, while a cache write costs roughly 125% of base (for the 5-minute TTL) or 200% (for the 1-hour TTL). If you misuse caching — putting the breakpoint on something that changes every request — you pay the premium every time and never get the discount. The feature has a foot-gun.

The short answer

prompt caching = KV cache for the prefix + persisted across requests + priced as a discount

Picture to keep: the model reading your 50,000-token prefix is a clerk pulling the same fat case file off the shelf and re-reading it cover to cover for every question you ask. Prompt caching is leaving the file open on the desk between questions — same file, same desk, and it gets cleared off if nobody comes back within a few minutes. Where the analogy breaks: a clerk would recognize a file that had one line edited. This matching is on the exact token prefix — change one character near the top and it’s a different file.

A request’s KV cache is normally thrown away when the request ends. Prompt caching keeps the cache for a marked prefix in GPU memory (or on a tier behind it) for some TTL — typically a few minutes — so a follow-up request that starts with the same prefix can skip prefill and start decoding almost immediately. The provider charges less because they did less work. That’s the entire idea.

How it works

Build it up from the naive version, one broken assumption at a time.

Naive attempt: keep the answer. The obvious “cache” is the one every web engineer reaches for — remember the response for a given input. Useless here. Your agent’s next request is not the same input; it’s the same 50,000 tokens plus a new tool result. And you want fresh generation anyway. Caching outputs saves nothing in an agent loop.

The fix: cache the intermediate state, not the answer. The KV cache for a prefix — the per-layer key and value tensors for tokens 1…N — is exactly the artifact that prefill produced. If the provider preserves it, the next request that begins with the same N tokens can attend against it directly and start generating. Prefill on N tokens is roughly O(N²) in the attention layers; reusing the cache means you skip recomputing the prefix entirely and only pay to process the new M-token suffix against it (roughly O(M·N + M²)). When M ≪ N — 200 new tokens against a 50,000-token prefix — that’s a huge saving, and the quadratic explosion in N is gone.

But how does the server know it’s the same prefix? It has to recognize your new request as an extension of a cached one, in the microseconds before it starts work, without re-reading 50,000 tokens. The open implementations do it by hashing: hash the prefix in chunks — typically token blocks — and look the hash up in a table of recently-computed prefixes belonging to your account. Match, and the cache entry is reused. Differ by one token, and the hash diverges from that point on. (vLLM’s open-source automatic prefix caching implements exactly this; the closed providers don’t publish their internals, so read the block-hash story as “how it’s done in the open implementations of the same idea.”)

That’s why “exact prefix match” is not a marketing simplification — it’s load-bearing. Insert a timestamp in your system prompt and every request hashes differently. Reorder your tool list and the cache misses. Swap a single token of whitespace and you’re paying full price.

Which forces the last constraint: the cache is not global. Caches are scoped per organization or workspace. Two different customers who happen to share a prefix do not share a cache — partly for isolation, and partly because the cached tensors are sitting on particular hardware, so a hit requires routing you back to where your state already lives. (The routing part is my inference from how these systems have to work, not something the provider docs spell out.)

What the API surface looks like

Providers expose the same underlying optimization differently:

The surface details differ; the underlying physics doesn’t. Somebody, somewhere, is keeping the prefix’s intermediate state around so prefill doesn’t have to run again.

What gets cached, what doesn’t

Across the major providers, the rule of thumb is: anything in the request that’s part of the prefix can be cached — system messages, tool definitions, message history, even images and documents in some cases. The output of the model is not “cached” in any useful sense; the same input still has to be re-decoded each time you ask for new tokens. (You can think of it this way: prefill is what gets reused, decode is what you pay for fresh.)

What breaks caching:

TTL and why caches expire fast

Five minutes feels short. My read on why it’s short: GPU memory is the most expensive memory in the data center, and a long-context KV cache is large — the exact per-request size depends on the model’s layer count, head dimensions, and cache precision, and providers don’t publish theirs, but for a 50,000-token prefix on a frontier-scale model it’s gigabytes, not megabytes. Keeping millions of users’ prefixes resident “just in case” would tile out the GPUs immediately. The TTL is a brutal eviction policy disguised as a feature. Anthropic’s 1-hour option costs 2× the write price; the docs don’t spell out the internal cost rationale, and both plausible explanations — pinning more VRAM, or pushing the cache to a slower tier — fit it equally well.

DeepSeek’s disk-backed approach is a different bet: cheaper to retain, slower to fetch. Whether disk- or VRAM-backed wins depends on the access pattern. For long-tail prefixes that get hit hours later, disk wins. For chatty agent loops where the next request lands in 2 seconds, VRAM wins because you can avoid the reload entirely.

Where it gets subtle

You started with prompt caching = KV cache for the prefix + persisted across requests + priced as a discount. What did this post add that the pricing page doesn’t? — + exact-prefix hashing. Prompt caching isn’t a feature stapled onto inference; it’s the KV cache that already makes generation fast within one request, given a longer lifetime so it can pay off across requests. And because the match is a hash of the prefix bytes, the discount is decided entirely by whether your 50,000 tokens are byte-identical to last turn’s — which is why a single timestamp at the top of that prefix flips you from Anthropic’s ~10%-of-base cache hit to a 125%-of-base cache write, every single turn.

Check yourself

Before you go — that same 50,000-token agent prefix, unchanged for weeks. You add one line at the very top: Current time: 2026-08-11T14:03:22Z, refreshed on every request. Your bill goes up, not down, even though you enabled caching. Walk through what happened.

Answer

Matching is on the prefix, from token 1 forward. A timestamp at the top changes the first tokens, so every request hashes differently and never hits — but the content is still marked cacheable, so every request performs a cache write, billed at a premium over base input (on Anthropic, 1.25× for the 5-minute TTL, 2× for the 1-hour; other providers price writes differently or not at all). You pay the write premium on 50,000 tokens forever and collect the read discount never. Moving the timestamp to the end of the prompt — after the cache breakpoint — fixes it, because everything before the breakpoint stays byte-identical.

And one more: two agents have identical token volumes, but agent A generates 200 output tokens per step and agent B generates 5,000. Caching helps which one more, and why?

Answer

Agent A. Prompt caching only touches prefill — the pass over the input. Decode still runs token by token at full price for every output token, cache or no cache. Agent B’s bill is dominated by output tokens that caching can’t touch, so a large discount on the input side moves a smaller share of its total. This is also why prompt caching stacks with, rather than competes with, decode-side tricks like speculative decoding.

Going deeper