Why retry with exponential backoff — and why jitter?
Retrying on failure sounds simple until you ship it at scale. Hammer the server and you make outages worse; back off but synchronize, and you accidentally rebuild the herd. Backoff is the timing rule; jitter is the part that keeps it from biting itself.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Naive attempt: retry now. Fix: wait, doubling each time
- But a shared ladder is still a synchronized ladder. Fix: jitter
- But a polite retry loop still never ends. Fix: a stopping rule
- The seams nobody puts in the diagram
- Famous related terms
- Going deeper
The picture version
Four pictures for a reader who has written a retry loop. The prose below fills in the seams the pictures skip.
1 · The problem
A two-second hiccup becomes a forty-minute outage.
2 · The half-fix
You replaced a stampede with a metronome.
3 · The fix
Each client picks a time from the window, not the edge of it.
4 · Keep this card
The whole thing on one index card.
Why it exists
An upstream API hiccups for two seconds. Your 500 worker processes all get a 503 in the same instant, all retry immediately, and the two-second hiccup becomes a forty-minute outage — the upstream now has 500 extra requests per loop iteration on top of the load it was already failing under. Nobody deployed anything. Nothing was “down” when it started. Those 500 workers are the example this post follows.
Retrying immediately is the naive move, and in a loop it’s a denial-of-service attack against a service that was already struggling. You arrive at the bad moment with friends.
So engineers learned a softer rule: wait, then retry, and wait longer each time you fail. This is exponential backoff — the wait doubles (or grows by some constant factor) on each failure. The intuition is that you don’t know how long the bad condition will last, and exponential growth covers a huge range of timescales — milliseconds to minutes — without you having to guess in advance.
That’s half the answer. The other half is the part that surprises people the first time they meet it: if all 500 workers retry on the same exponential schedule, you’ve replaced a stampede with a metronome. All 500 failed at second 0, so all 500 retry at second 1. They fail again, so all 500 come back at second 3. Then second 7. The server sees neat periodic spikes that, if anything, are easier to overload than random arrivals would be — because every spike is a near-simultaneous burst.
The fix is jitter — randomization on the wait. Each client picks its retry time from a window, not a point. The thundering herd flattens into something the server can actually serve.
Why it matters now
Anything that calls a service it doesn’t control needs a retry policy, and the places it’s already decided for you are worth knowing:
- Cloud SDKs. AWS, Google Cloud, and Azure clients commonly retry transient errors with backoff and jitter, but the defaults are SDK-specific (which errors qualify, what the cap is, what shape of jitter) and not always documented in the place you’d expect. Read the one you actually use.
- LLM provider APIs. Rate limits and 5xx are routine, especially during peak hours. Naive retry loops in agent code can trip the provider’s abuse limits and earn you a longer cooldown than the original error caused. Some official client libraries retry with backoff out of the box; many home-grown wrappers don’t.
- Webhook delivery. Stripe automatically retries failed webhook deliveries on an exponential schedule for up to three days. GitHub, by contrast, does not auto-redeliver — it surfaces failed deliveries and leaves the redelivery to you. The pattern shows up across webhook senders, but the policies vary; check the docs of the one you receive from.
- Distributed databases and queues. Leader elections, replica catch-up, and client connection retries use backoff. Without it, one node’s reboot can turn into a cluster-wide reconnect storm.
- TCP itself. TCP retransmits unacknowledged segments after a RTO; RFC 6298 specifies that on each retransmission-timer expiry the RTO doubles, with an implementation-defined upper bound. The pattern is so load-bearing the lower layers ship it whether you asked or not. The cost of getting it wrong is no longer “my script is annoying.” It’s “my agent fleet just synchronized into a wave that took down the upstream service for everyone.”
The short answer
retry policy = exponential backoff + jitter + a stopping rule
Picture to keep: 500 people who all got a busy signal at once — backoff tells them to wait longer between each redial, and jitter is what stops them from all redialling on the same tick of the clock. Where the picture breaks: real callers hear each other give up and can see the queue. Your 500 workers are blind to one another, which is why the synchronization has to be broken by randomness rather than by anyone noticing.
Three pieces. Exponential backoff spaces the retries out so a slow or flapping service has time to recover. Jitter breaks the synchronization between independent clients so their retries don’t pile up at the same instant. A stopping rule — max retries, max total wait, or both — keeps the loop from becoming an infinite background apology. Drop any of the three and you’ve built a footgun.
How it works
Each piece below exists because the previous one broke.
Naive attempt: retry now. Fix: wait, doubling each time
Pick a base delay (say 100 ms) and a factor (usually 2). The wait before
attempt n is base × factor^(n-1), capped at some ceiling so it doesn’t
drift into “retry next Tuesday” territory:
attempt 1: fail, wait 100 ms
attempt 2: fail, wait 200 ms
attempt 3: fail, wait 400 ms
attempt 4: fail, wait 800 ms
...
attempt k: fail, wait min(base × 2^(k-1), cap)
Why exponential and not linear? Linear backoff tends to spend most of the retry budget on small early waits and never reaches an interval long enough to ride out a real multi-second outage. Exponential covers many orders of magnitude with few retries: six attempts at base 100 ms reaches several seconds; ten reaches a minute or two. That’s the right shape for “I don’t know if this is a 50 ms blip or a 30 second deploy.”
But a shared ladder is still a synchronized ladder. Fix: jitter
Plain backoff schedules all 500 workers onto the same rungs, so the load arrives in spikes instead of continuously — and a spike is what overloads a server. Add jitter and each client picks a random delay from a window. The interesting question is which window.
The variants Marc Brooker used in the canonical AWS Architecture Blog post on this (March 4, 2015, “Exponential Backoff And Jitter”):
- No jitter —
wait = base × 2^(n-1). Simple, synchronizes clients, thundering-herd-prone. Don’t ship this. - Equal Jitter — half the wait is deterministic, half is random inside the current exponential window.
- Full Jitter —
wait = random(0, base × 2^(n-1)). Each retry lands uniformly somewhere inside the current exponential interval. - Decorrelated Jitter — the next wait is randomized within a window that grows from the previous wait, not from a fixed exponential schedule.
Brooker’s simulations on a synthetic workload found Full Jitter reduced server load substantially compared to no jitter. Against Decorrelated Jitter the two traded places depending on what you measure: in his words, Full Jitter “uses less work, but slightly more time.” The headline result — “spread the retries out randomly across the exponential window” — is the part that stuck, and it’s the only part worth memorising. Which variant a specific SDK ships varies by SDK and version, so read the one you depend on rather than assuming.
The shape worth memorizing: exponential window, uniform random within it. Pick a variant, document it, and move on.
But a polite retry loop still never ends. Fix: a stopping rule
Backoff without a stopping rule is a quiet disaster. Two common shapes:
- Bounded retries — give up after N attempts, surface the error to the caller. Right for interactive paths where someone is waiting.
- Bounded total time — keep retrying until a deadline (e.g. the caller’s overall timeout) passes. Right for background work where the job has a wall-clock budget anyway.
Both have a subtler companion: the per-attempt timeout has to be shorter than the deadline. A retry loop where each attempt blocks for 30 seconds with a 30-second total budget gives you exactly one try. The classic “we have retries, why didn’t they help?” postmortem ends here more often than it should.
The seams nobody puts in the diagram
- Don’t retry non-idempotent writes blindly. A
POST /chargesthat times out may have succeeded; retrying creates a duplicate charge. Use idempotency keys, or restrict retries to GETs and explicitly safe endpoints. The HTTP method is a hint, not a guarantee. - Honor
Retry-After. When a server returns 429 or 503 with aRetry-Afterheader, that value is the server telling you exactly how long to wait. Your local backoff math should defer to it, not fight it. - Circuit breakers complement backoff. Backoff handles a single call. A circuit breaker handles the aggregate — when failure rate crosses a threshold, stop trying for a window. Without one, every caller in your fleet keeps generously retrying a service that is plainly down.
- Retry budgets stop runaway amplification. A 3-retry policy turns one failure into four requests. Cascade that across three services and you’ve quietly made each user-facing failure cost 64 backend requests. Some systems cap retries at the fleet level (e.g. “no more than 10% of total traffic may be retries”) to bound this.
- Jitter doesn’t help one client. A single client retrying a single
call gains nothing from jitter; the herd-flattening effect only shows
up across many independent clients. The cost of adding it is small —
one extra
random()per retry — and the moment your one client becomes ten, you’re glad you did.
You started with retry policy = exponential backoff + jitter + a stopping rule. What did the 500 workers add? — + backoff is about the fleet, not about your call. Every part of the policy is chosen for what happens when
many independent clients hit the same bad moment; measured from a single
client, backoff just looks like an artificial delay, which is exactly why it
keeps getting stripped out of home-grown retry loops.
Famous related terms
- Idempotency —
idempotent op = same result whether you run it once or N times— the property that makes retrying a write safe. Without it, retries are guesses about a write you may have already done. - Circuit breaker —
circuit breaker = error counter + cooldown window— short-circuits calls to a downstream that’s clearly down, so you stop wasting capacity on doomed retries. - Token bucket / leaky bucket —
token bucket = bucket of N tokens + refill rate— a rate-limiting shape on the server side that backoff is reacting to. Knowing the bucket helps you size the backoff. - Thundering herd —
thundering herd = many clients waking at once + one shared resource— the failure mode jitter is built to kill. - TCP congestion control — backoff’s older cousin. TCP halves its sending rate on loss and grows it back slowly; same instinct, different layer.
- Hedged requests —
hedged request = send the call + send a duplicate after a small delay if the first hasn't returned— the opposite of backoff for tail-latency-sensitive reads. They live together in the same toolbox.
Going deeper
- Marc Brooker, “Exponential Backoff And Jitter” (AWS Architecture Blog, 2015) — the primary source for the question “which jitter shape actually wins,” with the simulations behind the answer.
- Google’s SRE Book, “Handling Overload” — read this for the fleet-level question: what does retry amplification do to a system that’s already saturated, and what bounds it?
- RFC 9110 §10.2.3
(
Retry-After) — the rabbit hole for “what is the server allowed to tell me about when to come back,” which your backoff math should defer to.