Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why retry with exponential backoff — and why jitter?

Retrying on failure sounds simple until you ship it at scale. Hammer the server and you make outages worse; back off but synchronize, and you accidentally rebuild the herd. Backoff is the timing rule; jitter is the part that keeps it from biting itself.

Networking intro Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Four pictures for a reader who has written a retry loop. The prose below fills in the seams the pictures skip.

1 · The problem

A two-second hiccup becomes a forty-minute outage.

500 workers all retry at once an upstream that was already failing under the load it had now +500 requests per loop iteration on top of that Nobody deployed anything. Nothing was down when it started. retrying immediately is, in a loop, a denial-of-service attack against a service that was already struggling You arrive at the bad moment with friends.
Those 500 workers are the running example for the whole post. The obvious softer rule is: wait, then retry, and wait longer each time you fail — which covers milliseconds to minutes without your having to guess in advance how long the bad condition lasts.

2 · The half-fix

You replaced a stampede with a metronome.

t = 0 all 500 all 500 all 500 all 500 1s 3s 7s The waits doubled. The synchronisation didn’t break. every client failed at the same instant, so every client is on the same rung of the same ladder and a near-simultaneous burst is, if anything, easier to overload with than random arrivals would be
Exponential backoff on its own solves the rate problem and leaves the timing problem untouched. The clients are blind to one another, which is exactly why the fix has to come from randomness rather than from anyone noticing the queue.

3 · The fix

Each client picks a time from the window, not the edge of it.

the exponential window is the same; where inside it you land is now random window 1 window 2 Same total retries. Spread continuously instead of stacked on one tick. Full Jitter picks uniformly from zero to the whole window; Equal Jitter keeps half the wait fixed and randomises the rest. Brooker’s comparison: Full Jitter “uses less work, but slightly more time” than Decorrelated Jitter — the two trade places depending on what you measure.
The variant matters less than people expect. The shape worth memorising is: exponential window, uniform random inside it — pick a variant, document which one, and move on.

4 · Keep this card

The whole thing on one index card.

retry policy = exponential backoff — so a flapping service can recover + jitter — so clients stop arriving together + a stopping rule — so it ends ∴ drop any one of the three and you have built a footgun
Picture to keep: 500 people who all got a busy signal at once — backoff tells them to wait longer between each redial, and jitter is what stops them all redialling on the same tick of the clock. Where the picture breaks: real callers hear each other give up and can see the queue. Your 500 workers are blind to one another.

Why it exists

An upstream API hiccups for two seconds. Your 500 worker processes all get a 503 in the same instant, all retry immediately, and the two-second hiccup becomes a forty-minute outage — the upstream now has 500 extra requests per loop iteration on top of the load it was already failing under. Nobody deployed anything. Nothing was “down” when it started. Those 500 workers are the example this post follows.

Retrying immediately is the naive move, and in a loop it’s a denial-of-service attack against a service that was already struggling. You arrive at the bad moment with friends.

So engineers learned a softer rule: wait, then retry, and wait longer each time you fail. This is exponential backoff — the wait doubles (or grows by some constant factor) on each failure. The intuition is that you don’t know how long the bad condition will last, and exponential growth covers a huge range of timescales — milliseconds to minutes — without you having to guess in advance.

That’s half the answer. The other half is the part that surprises people the first time they meet it: if all 500 workers retry on the same exponential schedule, you’ve replaced a stampede with a metronome. All 500 failed at second 0, so all 500 retry at second 1. They fail again, so all 500 come back at second 3. Then second 7. The server sees neat periodic spikes that, if anything, are easier to overload than random arrivals would be — because every spike is a near-simultaneous burst.

The fix is jitter — randomization on the wait. Each client picks its retry time from a window, not a point. The thundering herd flattens into something the server can actually serve.

Why it matters now

Anything that calls a service it doesn’t control needs a retry policy, and the places it’s already decided for you are worth knowing:

The short answer

retry policy = exponential backoff + jitter + a stopping rule

Picture to keep: 500 people who all got a busy signal at once — backoff tells them to wait longer between each redial, and jitter is what stops them from all redialling on the same tick of the clock. Where the picture breaks: real callers hear each other give up and can see the queue. Your 500 workers are blind to one another, which is why the synchronization has to be broken by randomness rather than by anyone noticing.

Three pieces. Exponential backoff spaces the retries out so a slow or flapping service has time to recover. Jitter breaks the synchronization between independent clients so their retries don’t pile up at the same instant. A stopping rule — max retries, max total wait, or both — keeps the loop from becoming an infinite background apology. Drop any of the three and you’ve built a footgun.

How it works

Each piece below exists because the previous one broke.

Naive attempt: retry now. Fix: wait, doubling each time

Pick a base delay (say 100 ms) and a factor (usually 2). The wait before attempt n is base × factor^(n-1), capped at some ceiling so it doesn’t drift into “retry next Tuesday” territory:

attempt 1: fail, wait 100 ms
attempt 2: fail, wait 200 ms
attempt 3: fail, wait 400 ms
attempt 4: fail, wait 800 ms
...
attempt k: fail, wait min(base × 2^(k-1), cap)

Why exponential and not linear? Linear backoff tends to spend most of the retry budget on small early waits and never reaches an interval long enough to ride out a real multi-second outage. Exponential covers many orders of magnitude with few retries: six attempts at base 100 ms reaches several seconds; ten reaches a minute or two. That’s the right shape for “I don’t know if this is a 50 ms blip or a 30 second deploy.”

But a shared ladder is still a synchronized ladder. Fix: jitter

Plain backoff schedules all 500 workers onto the same rungs, so the load arrives in spikes instead of continuously — and a spike is what overloads a server. Add jitter and each client picks a random delay from a window. The interesting question is which window.

The variants Marc Brooker used in the canonical AWS Architecture Blog post on this (March 4, 2015, “Exponential Backoff And Jitter”):

Brooker’s simulations on a synthetic workload found Full Jitter reduced server load substantially compared to no jitter. Against Decorrelated Jitter the two traded places depending on what you measure: in his words, Full Jitter “uses less work, but slightly more time.” The headline result — “spread the retries out randomly across the exponential window” — is the part that stuck, and it’s the only part worth memorising. Which variant a specific SDK ships varies by SDK and version, so read the one you depend on rather than assuming.

The shape worth memorizing: exponential window, uniform random within it. Pick a variant, document it, and move on.

But a polite retry loop still never ends. Fix: a stopping rule

Backoff without a stopping rule is a quiet disaster. Two common shapes:

Both have a subtler companion: the per-attempt timeout has to be shorter than the deadline. A retry loop where each attempt blocks for 30 seconds with a 30-second total budget gives you exactly one try. The classic “we have retries, why didn’t they help?” postmortem ends here more often than it should.

The seams nobody puts in the diagram

You started with retry policy = exponential backoff + jitter + a stopping rule. What did the 500 workers add? — + backoff is about the fleet, not about your call. Every part of the policy is chosen for what happens when many independent clients hit the same bad moment; measured from a single client, backoff just looks like an artificial delay, which is exactly why it keeps getting stripped out of home-grown retry loops.

Going deeper