Why idempotency keys exist
The network can drop your response after the work is done. Now you have to retry — and you have no idea whether you'd be doing it for the first time or the second. Idempotency keys are the small protocol the client and server agree on so the retry is safe.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: let the server dedupe by content
- Attempt 2: let the server hand out a token first
- Attempt 3: the client mints the key before it sends anything
- Attempt 4: store the key next to the work
- Attempt 5: make sure the same key means the same request
- Attempt 6: let keys expire — but not too soon
- Show the seams
- Famous related terms
- Going deeper
The picture version
Five pictures for a reader who has never had to think about what a timeout actually means. The prose below fills in the seams the pictures skip.
1 · The problem
The reply got lost. The money may not have.
2 · Two fixes that don’t work
The server can’t tell a retry from a second purchase.
3 · The move
Bring your own ticket, and bring it every time.
4 · The part that gets built wrong
If the key and the work can commit apart, they will.
5 · Keep this card
The whole thing on one index card.
Why it exists
You call POST /charges for $42. The request goes out. The connection times
out. Did the charge happen?
You don’t know. The packet that would have told you got lost — but the packet that asked for the work might have arrived just fine. From the client’s seat, “request succeeded but reply was dropped” is indistinguishable from “request never made it.” Both look like a timeout.
So you have to choose, and both choices are bad:
- Retry. You might charge the customer twice.
- Don’t retry. You might never charge them, and your code will think the operation succeeded silently or fail loudly.
This is the core problem. The network can hide the answer without hiding the work. Before reading on: what could the client put in that second request that would let the server tell “this is the retry of the $42 charge you already ran” apart from “here is a fresh $42 charge”?
Idempotency
is the way out: design the operation so retrying it is harmless. For a
GET that’s free — reading the same row twice is the same as reading it
once. For a write, you usually need help. The help is an
idempotency key: a token the client invents and sends along, that
the server uses to recognize “I’ve already done this exact request,
here’s the same answer I gave last time.”
The key turns a dangerous retry into a safe lookup. We’ll follow that same $42 charge the rest of the way down.
Why it matters now
Anywhere the cost of a duplicate is real, idempotency keys appear:
- Payments. Stripe’s v1 API takes an
Idempotency-Keyheader on everyPOST; resending the same key returns the originally saved response instead of creating a second charge or refund. (Stripe’s newer v2 API extends keys toDELETEand gives them a much longer replay window — read the docs of the version you’re calling.) The header is the canonical example most engineers meet first. - LLM and other expensive AI calls. A retried request that re-runs the model costs you the dollars and the latency a second time, even if the first one actually completed. Whether a given provider dedupes a retried request is a per-API question with no default answer — some document an idempotency key, others document nothing — and a home-grown agent retry loop firing on a timeout has no way to know whether the first call completed. Check the docs of the API you’re calling before assuming a retry is free.
- Webhook receivers. Senders that retry — Stripe automatically redelivers failed webhooks for up to three days in live mode; many others have similar policies (GitHub, by contrast, does not automatically redeliver — failed deliveries have to be replayed deliberately, from the UI or the API) — can deliver the same event more than once. A webhook handler that isn’t idempotent will, eventually, ship a duplicate side effect to production. The remedy is the same shape: dedupe by the event ID the sender provides.
- Message queues. Standard SQS queues and Kafka’s default consumer semantics are at-least-once, not exactly-once. (RabbitMQ’s guarantees depend on which acknowledgement mode the consumer picks: automatic acknowledgement counts a message as delivered the moment it is sent, so a consumer crash loses it, while manual acknowledgement is what lets the broker redeliver.) Consumers have to dedupe themselves; an idempotency key on the message is how.
- Background job systems. A worker that crashes after doing the work but before acknowledging the job will see the same job again on restart.
The pattern shows up wherever a network or process boundary can swallow an acknowledgement. Which is, on a long enough timeline, everywhere.
The short answer
idempotency key = client-generated unique ID + server-side dedupe table
Picture to keep: a coat check. You hand over the coat and keep the numbered ticket. Come back with the same ticket and you get the same coat back — not a second coat, and not an argument about whether you were here before. Except that in a coat check they print the ticket; here you have to bring your own, because the moment they hand you one is exactly the moment that can get lost.
The client picks a unique value (often a UUID) before it sends the request, and reuses that same value for every retry of that logical operation. The server, before doing the work, checks a dedupe store: “have I seen this key?” If yes, return the recorded response. If no, do the work, record the response under the key, then return it. Two halves: the client’s promise that the same key means the same intent, and the server’s commitment to remember.
How it works
Attempt 1: let the server dedupe by content
The obvious move needs nothing from the client. The server hashes the request body — amount, currency, customer — and refuses anything it has seen recently.
Why it breaks: it can’t tell a retry from a repeat. A customer who genuinely buys two $42 things a minute apart sends two byte-identical requests, and the second one silently vanishes. Content identifies the shape of the request; it can’t identify the intent.
Attempt 2: let the server hand out a token first
So make the client ask: POST /charge-tokens returns a fresh token, then
the client spends that token on the real request.
Why it breaks: you’ve moved the problem, not solved it. The token response can be lost in exactly the same way the charge response was, and now the client doesn’t know whether it has a token. You’d need an idempotency mechanism for the token endpoint — which is where the recursion should tip you off.
Attempt 3: the client mints the key before it sends anything
That’s the fix, and the reason it has to be this way is worth sitting with. The whole point is to survive the case where the server’s response never arrives — so the client must be able to name the operation before it has any reply to anchor on. A server-generated identifier is fine for referencing a created resource afterwards; it can’t deduplicate the request that created it.
For our $42 charge, with key k:
- Client generates
konce, before the first attempt. SendsIdempotency-Key: kplus the request body. - Server looks up
kin its dedupe store.- Hit, with a stored response → return it. The work has already been done; the client just didn’t hear about it.
- Miss → mark
kas in-progress, do the work, record the response underk, return the response. - Hit, but still in-progress → either wait for the first attempt to finish, or return a “duplicate-in-flight” error so the client backs off and retries later.
- Client retries on a timeout or a
5xx
server error, with the same
k, and gets either the original outcome or a quick deterministic error.
That’s the mechanism. Every remaining refinement is a hole someone fell into.
Attempt 4: store the key next to the work
A dedupe store is at minimum a map from key to “I’m working on it” or “here’s the recorded result.” Keep it in Redis, say, and write the charge to Postgres.
Why it breaks: the two can disagree. Crash between “did the work” and
“recorded key k” and the retry does the work again; crash the other way
and work that never happened looks done. The fix is atomicity: put the key
in a row of the same database the work writes to, and commit both in one
transaction. This is the part most home-grown implementations get wrong.
That fix is complete only when the work is a database write. Our $42
charge isn’t — the card gets charged by a processor across the network,
which your transaction can’t roll back. There the honest version is: your
transaction covers the key row and your own records, and the external call
needs its own idempotency key sent downstream (which is exactly why
Stripe offers one). You don’t escape the problem, you push the deduping to
the boundary that owns the side effect. The outbox pattern in the seams
below is the general form of this.
A common shape:
INSERT INTO idempotency_keys (key, request_fingerprint, status)
VALUES ($1, $2, 'in_progress')
ON CONFLICT (key) DO NOTHING;
If the insert wins, this attempt does the work and updates the row to
completed with the response, in the same transaction as the business
write. If it loses, the row already exists — and note that the INSERT
alone doesn’t tell you what is there. The server still has to read the
row back and branch: a completed row means hand over the stored response;
an in_progress row means the first attempt is still running, which is the
awkward case covered below.
Attempt 5: make sure the same key means the same request
Idempotency keys are a promise about intent, not just identity. If
the client sends key k once with {amount: 42} and again with
{amount: 4200}, the server should not silently treat the second as a
duplicate of the first — that would let a bug or a race quietly
overwrite a charge with a different one.
The defensive move is for the server to also store a fingerprint of the request body and reject mismatches. Stripe’s API documents exactly this behavior — replaying a key with a different payload is an error, not a silent dedupe. Anything less is a footgun.
Attempt 6: let keys expire — but not too soon
Storing every key forever is unbounded growth. Most real systems give keys a TTL — Stripe’s v1 docs say keys can be pruned once they’re at least 24 hours old; v2 extends the replay window much further. The retention has to comfortably exceed the client’s worst-case retry budget, or else the client retries with a key the server has already forgotten, the dedupe miss looks fresh, and the operation runs again.
Show the seams
- Idempotent in HTTP’s sense isn’t quite the same thing. RFC 9110
calls a method “idempotent” when N identical requests have the same
effect as one.
PUTandDELETEare defined that way;POSTis not. But that’s a property of the method, not of any particular request. Idempotency keys layer “this specific request is a retry of that one” on top, so they makePOSToperationally idempotent for a given key. - Concurrent retries are the awkward case. A client that sends
request
ktwice in parallel (because the first attempt hasn’t obviously failed yet) puts the server in the in-progress state twice. Servers handle this by either serializing on the row lock or by returning a 409-ish “operation in progress, retry later.” Either is fine; silently doing the work twice is not. - Exactly-once is still a fairy tale at the network layer. What idempotency keys give you is effectively-once application semantics on top of at-least-once delivery. The duplicates still arrive; the server just refuses to act on them twice. If anyone promises you exactly-once delivery, ask where the dedupe lives — it’s always somewhere.
- The key has to be generated before the first send. If a client generates the key inside its retry loop, every retry gets a fresh key and the dedupe never fires. This sounds obvious and is one of the most common bugs.
- Side effects outside the database don’t roll back. If the work sends an email and then fails to record the idempotency row, a retry will send a second email. The general pattern — record the side effect’s intent transactionally, perform the side effect from an outbox — is the outbox pattern, and idempotency keys live happily next to it.
You started with idempotency key = client-generated unique ID + server-side dedupe table. What did the $42 charge force into that line? —
+ the key and the work committing in one transaction. Everything else
here (fingerprints, in-progress rows, TTLs) hardens the contract, but if
the key and the charge can commit separately, the dedupe table is just a
log of things you think happened.
Famous related terms
- UUID —
UUID = 128 bits + a versioned generation rule that makes collisions negligible without a coordinator— the usual choice for the key itself, because the client can mint one alone. - At-least-once delivery —
at-least-once = retry on uncertainty + accept the cost of duplicates— the delivery mode idempotency keys exist to compensate for. - Outbox pattern —
outbox = side-effect intent table + drain worker— pairs with idempotency keys when the work has external side effects. - Exponential backoff —
retry policy = backoff + jitter + stopping rule— the when of retrying; idempotency keys are the how to make it safe. - Two-phase commit —
2PC = prepare phase + commit phase across participants— the heavyweight alternative when you can’t tolerate even a transient duplicate. Almost always the wrong tool for an HTTP API. - CAS —
CAS = read expected value + atomic conditional write— the in-process cousin; same instinct (“only do this if the world hasn’t moved”), different scale.
Going deeper
- Stripe’s API reference on the
Idempotency-Keyheader — the primary source for the exact contract a widely-copied production implementation commits to, including what happens when you replay a key with a changed body. - Brandur Leach’s “Designing robust and predictable APIs with idempotency” on the Stripe engineering blog — the explainer for “how do I actually build the server side of this,” transaction boundaries and all.
- Tyler Treat, “You Cannot Have Exactly-Once Delivery” — the rabbit hole, for why effectively-once on top of at-least-once is the honest target and exactly-once delivery isn’t on the menu.