Why does prompt caching exist?
Your agent sends the same 50,000-token system prompt on every turn. Providers charge a fraction of the usual rate when they recognize it — not out of generosity, but because they stopped doing the work.
On this page
The picture version
Six pictures for a reader who has never looked at an inference bill. The prose below fills in the seams the pictures skip.
1 · The problem
99% of what you send is what you already sent.
2 · The wrong cache
Caching the answer saves nothing here.
3 · How the server recognises you
It hashes the prefix in blocks. One token differs, everything after diverges.
4 · The foot-gun
Miss every time and the bill goes up.
5 · Why it expires so fast
Five minutes is an eviction policy wearing a feature’s clothes.
6 · Keep this card
The whole thing on one index card.
Why it exists
Build an agent for five minutes and you’ll notice something uncomfortable.
Every turn, you’re sending the model the same long system prompt, the same tool definitions, the same scrollback of “here’s what we did last step.” The variable part — the user’s new message, or the result of the last tool call — is tiny. Maybe 1% of the input tokens. The other 99% is identical to what you sent on the previous request.
The provider, on the other hand, is in an even more uncomfortable position. From their side: a request comes in with 50,000 tokens of prefix. Their server runs the full prefill pass — quadratic in prompt length, the most expensive thing the model does — to populate the KV cache for those 50,000 tokens. The model emits a few hundred tokens of output. The request ends. The KV cache is freed. Three seconds later, the same client sends 50,001 tokens — the original 50,000 plus one new user message. The server does the entire 50,000-token prefill again, from scratch, because nothing was kept around.
That is an absurd amount of duplicated work. Both sides know it. Prompt caching is the obvious move: keep the KV cache for the prefix around between requests (in GPU memory, or on a slower tier behind it), and bill the user for the cache hit instead of the recomputation. Prefill is the most expensive part of inference for long prompts. Skipping it is what pays for the discount on the pricing page — on Anthropic’s API, a cache hit bills at roughly 10% of the base input rate. (Discounts differ by provider; the shape doesn’t.)
Why it matters now
A few years ago, prompts were short. “Summarize this paragraph.” Caching across requests would have saved nothing worth the engineering. Three things changed:
- Long system prompts. Modern assistants and coding agents ship with 5,000–50,000 tokens of system instructions, persona, formatting rules, and examples before the user has typed a single character.
- Long tool / MCP definitions. A capable agent harness registers dozens of tools. Their JSON schemas alone are often thousands of tokens. They almost never change between turns.
- Multi-turn loops. An agent that runs ten tool calls to answer one question sends the entire prior trace as input on each step. Turn N’s input is turn N-1’s input plus a few hundred tokens.
In all three cases, the prefix is stable and the suffix is small. That’s the exact regime where caching wins. Without it, agentic workloads would be priced as if every step is a fresh request — which is to say, mostly priced as prefill cost on a prompt that was already prefilled five seconds ago.
It also matters because the cost model is not symmetric. On Anthropic’s API, a cache hit is priced at roughly 10% of the base input rate, while a cache write costs roughly 125% of base (for the 5-minute TTL) or 200% (for the 1-hour TTL). If you misuse caching — putting the breakpoint on something that changes every request — you pay the premium every time and never get the discount. The feature has a foot-gun.
The short answer
prompt caching = KV cache for the prefix + persisted across requests + priced as a discount
Picture to keep: the model reading your 50,000-token prefix is a clerk pulling the same fat case file off the shelf and re-reading it cover to cover for every question you ask. Prompt caching is leaving the file open on the desk between questions — same file, same desk, and it gets cleared off if nobody comes back within a few minutes. Where the analogy breaks: a clerk would recognize a file that had one line edited. This matching is on the exact token prefix — change one character near the top and it’s a different file.
A request’s KV cache is normally thrown away when the request ends. Prompt caching keeps the cache for a marked prefix in GPU memory (or on a tier behind it) for some TTL — typically a few minutes — so a follow-up request that starts with the same prefix can skip prefill and start decoding almost immediately. The provider charges less because they did less work. That’s the entire idea.
How it works
Build it up from the naive version, one broken assumption at a time.
Naive attempt: keep the answer. The obvious “cache” is the one every web engineer reaches for — remember the response for a given input. Useless here. Your agent’s next request is not the same input; it’s the same 50,000 tokens plus a new tool result. And you want fresh generation anyway. Caching outputs saves nothing in an agent loop.
The fix: cache the intermediate state, not the answer. The KV cache for a prefix — the per-layer key and value tensors for tokens 1…N — is exactly the artifact that prefill produced. If the provider preserves it, the next request that begins with the same N tokens can attend against it directly and start generating. Prefill on N tokens is roughly O(N²) in the attention layers; reusing the cache means you skip recomputing the prefix entirely and only pay to process the new M-token suffix against it (roughly O(M·N + M²)). When M ≪ N — 200 new tokens against a 50,000-token prefix — that’s a huge saving, and the quadratic explosion in N is gone.
But how does the server know it’s the same prefix? It has to recognize your new request as an extension of a cached one, in the microseconds before it starts work, without re-reading 50,000 tokens. The open implementations do it by hashing: hash the prefix in chunks — typically token blocks — and look the hash up in a table of recently-computed prefixes belonging to your account. Match, and the cache entry is reused. Differ by one token, and the hash diverges from that point on. (vLLM’s open-source automatic prefix caching implements exactly this; the closed providers don’t publish their internals, so read the block-hash story as “how it’s done in the open implementations of the same idea.”)
That’s why “exact prefix match” is not a marketing simplification — it’s load-bearing. Insert a timestamp in your system prompt and every request hashes differently. Reorder your tool list and the cache misses. Swap a single token of whitespace and you’re paying full price.
Which forces the last constraint: the cache is not global. Caches are scoped per organization or workspace. Two different customers who happen to share a prefix do not share a cache — partly for isolation, and partly because the cached tensors are sitting on particular hardware, so a hit requires routing you back to where your state already lives. (The routing part is my inference from how these systems have to work, not something the provider docs spell out.)
What the API surface looks like
Providers expose the same underlying optimization differently:
- Anthropic (Claude) — supports both automatic caching (a single top-level
cache_controlfield that moves forward as the conversation grows) and explicit block-level breakpoints (cache_control: { type: "ephemeral" }on individual content blocks, up to 4 per request). Default TTL 5 minutes; an extended 1-hour TTL is available at a higher write cost. There’s a minimum cacheable prefix length that varies by model — small models need more tokens before caching kicks in — so check the docs for yours rather than assuming. Cache reads cost ~10% of base input; cache writes cost 1.25× (5m) or 2× (1h) of base input. - OpenAI — automatic at launch, no opt-in required (announced October 1, 2024); the API has since grown explicit cache controls, so check the current reference rather than assuming it is purely automatic. Prefixes of ≥1024 tokens are eligible; cache hits land in 128-token increments. The launch announcement gave a 50% discount on cached input tokens; current pricing is model-specific and varies, so check the model’s pricing page rather than trusting “50%” as a stable rule.
- DeepSeek — automatic, disk-backed (not in-VRAM). At the August 2024 launch, cache hits were priced at 1/10 of the base input rate. Current ratios vary per model. The disk-tier choice is a different point on the cost-vs-latency tradeoff: cheaper to keep around for longer, slower to load back.
The surface details differ; the underlying physics doesn’t. Somebody, somewhere, is keeping the prefix’s intermediate state around so prefill doesn’t have to run again.
What gets cached, what doesn’t
Across the major providers, the rule of thumb is: anything in the request that’s part of the prefix can be cached — system messages, tool definitions, message history, even images and documents in some cases. The output of the model is not “cached” in any useful sense; the same input still has to be re-decoded each time you ask for new tokens. (You can think of it this way: prefill is what gets reused, decode is what you pay for fresh.)
What breaks caching:
- Anything that changes the prefix bytes. Timestamps, request IDs, randomized examples, anything dated.
- Anything earlier in the hierarchy changing. Most providers cascade invalidation: if your tools change, the message caches downstream of them also invalidate, even if those bytes didn’t change. That isn’t an arbitrary policy — a token’s cached keys and values are computed by attending over everything before it, so change token 5 and every entry after it is describing a different context.
- Model swaps or parameter changes. Different weights produce different keys and values; the cache is model-specific. Some non-obvious parameters (e.g. tool-choice mode, certain feature flags) can invalidate parts of the cache too.
TTL and why caches expire fast
Five minutes feels short. My read on why it’s short: GPU memory is the most expensive memory in the data center, and a long-context KV cache is large — the exact per-request size depends on the model’s layer count, head dimensions, and cache precision, and providers don’t publish theirs, but for a 50,000-token prefix on a frontier-scale model it’s gigabytes, not megabytes. Keeping millions of users’ prefixes resident “just in case” would tile out the GPUs immediately. The TTL is a brutal eviction policy disguised as a feature. Anthropic’s 1-hour option costs 2× the write price; the docs don’t spell out the internal cost rationale, and both plausible explanations — pinning more VRAM, or pushing the cache to a slower tier — fit it equally well.
DeepSeek’s disk-backed approach is a different bet: cheaper to retain, slower to fetch. Whether disk- or VRAM-backed wins depends on the access pattern. For long-tail prefixes that get hit hours later, disk wins. For chatty agent loops where the next request lands in 2 seconds, VRAM wins because you can avoid the reload entirely.
Where it gets subtle
- Concurrent requests can race. If two requests with the same fresh prefix arrive simultaneously, both will trigger a cache write before either has finished. Anthropic’s docs explicitly recommend serializing the first request to populate the cache before fanning out, which is a tell that this case is real and annoying.
- The breakpoint must sit on stable content. If you put your
cache_controlmark on the dynamic suffix, the prefix hash includes the changing bytes and you cache-write a fresh entry every request — paying the write premium and never getting a hit. Mark the end of the static prefix instead. - The “discount” is also a sales argument for longer prompts. Once caching is on, the marginal cost of stuffing more examples or more docs into your system prompt falls by whatever the read discount is — an order of magnitude, on providers that price hits at ~10% of base. It’s a reasonable read — though not one anybody has measured — that this is part of why agent system prompts ballooned over 2024–2025: caching made size much less expensive than it used to be. Whether that’s a good thing for prompt quality is a separate question.
- Output-token cost is not affected. Caching only touches input pricing. If your agent generates lots of tokens, prompt caching helps less than you think.
- The numbers are a snapshot. The concrete pricing ratios and minimum-token thresholds here come from each provider’s public docs as of early 2026, and they drift. The shape of the cost (read ≪ base ≪ write, with write only worth it if you’ll hit it more than ~once) is the durable part; treat the specific multipliers as a snapshot.
You started with prompt caching = KV cache for the prefix + persisted across requests + priced as a discount. What did this post add that the pricing page doesn’t? — + exact-prefix hashing. Prompt caching isn’t a feature stapled onto inference; it’s the KV cache that already makes generation fast within one request, given a longer lifetime so it can pay off across requests. And because the match is a hash of the prefix bytes, the discount is decided entirely by whether your 50,000 tokens are byte-identical to last turn’s — which is why a single timestamp at the top of that prefix flips you from Anthropic’s ~10%-of-base cache hit to a 125%-of-base cache write, every single turn.
Check yourself
Before you go — that same 50,000-token agent prefix, unchanged for weeks. You add one line at the very top: Current time: 2026-08-11T14:03:22Z, refreshed on every request. Your bill goes up, not down, even though you enabled caching. Walk through what happened.
Answer
Matching is on the prefix, from token 1 forward. A timestamp at the top changes the first tokens, so every request hashes differently and never hits — but the content is still marked cacheable, so every request performs a cache write, billed at a premium over base input (on Anthropic, 1.25× for the 5-minute TTL, 2× for the 1-hour; other providers price writes differently or not at all). You pay the write premium on 50,000 tokens forever and collect the read discount never. Moving the timestamp to the end of the prompt — after the cache breakpoint — fixes it, because everything before the breakpoint stays byte-identical.
And one more: two agents have identical token volumes, but agent A generates 200 output tokens per step and agent B generates 5,000. Caching helps which one more, and why?
Answer
Agent A. Prompt caching only touches prefill — the pass over the input. Decode still runs token by token at full price for every output token, cache or no cache. Agent B’s bill is dominated by output tokens that caching can’t touch, so a large discount on the input side moves a smaller share of its total. This is also why prompt caching stacks with, rather than competes with, decode-side tricks like speculative decoding.
Famous related terms
- KV cache —
KV cache = per-layer K and V tensors for past tokens, reused on the next decode step. Prompt caching is this, persisted past the end of a request. - Prefill vs. decode —
prefill = one quadratic pass over the prompt; decode = cheap one-token-at-a-time steps. Prompt caching specifically targets prefill — the expensive phase. It does nothing for decode. - Cache breakpoint —
breakpoint = a marker on a content block saying "the prefix up to here is cacheable". Anthropic’s explicit form. OpenAI’s automatic mode is roughly “the breakpoint is the longest matched prefix, found for you.” - Cache hit rate —
hit rate = fraction of input tokens served from cache. The metric to actually optimize. Reported under provider-specific names —cache_read_input_tokens(Anthropic),cached_tokens(OpenAI),prompt_cache_hit_tokens(DeepSeek). - Continuous batching —
continuous batching = scheduling new tokens from many requests together each step. Orthogonal optimization on the decode side; stacks with prompt caching on the prefill side. - Speculative decoding —
speculative decoding = small model proposes, big model verifies in parallel. Another decode-side trick. Stacks with prompt caching. - Context window —
context window = max tokens the attention mechanism can address. Prompt caching makes large windows affordable; it doesn’t make them larger.
Going deeper
- Anthropic prompt caching docs — go here for the exact rules this post generalizes over: where breakpoints are allowed, what invalidates a cache, and the current read/write multipliers (the numbers above are a snapshot and will drift).
- vLLM — Automatic Prefix Caching design doc — read this if you want to see the hashing and block-reuse machinery spelled out in an open implementation, since no commercial provider documents its internals.
- Kwon et al., PagedAttention (vLLM, 2023) — the rabbit hole: how KV-cache memory is actually managed inside a serving engine, and why block-level storage is what makes cross-request prefix reuse practical at all.