Why circuit breakers exist
Backoff makes a single retry polite. But when a downstream is plainly down, every caller in your fleet generously retrying it is the actual problem. A circuit breaker is the small piece that says: stop calling for a while — the answer isn't going to change in the next 50 ms.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: just stop calling
- Attempt 2: add a cooldown
- Attempt 3: probe before committing
- Attempt 4: count the right things
- Attempt 5: give each destination its own breaker
- Fallbacks: the part the diagram doesn’t show
- Show the seams
- Famous related terms
- Going deeper
The picture version
Four pictures for a reader who has already added retries and still went down. The prose below fills in the seams the pictures skip.
1 · The problem
Ten thousand polite callers are still a flood.
2 · The move
Put something in front of the call that can stop asking.
3 · The three states
And the middle one is the whole design.
4 · Keep this card
The whole thing on one index card.
Why it exists
Your checkout service calls a payments service on every order. You’ve already shipped retries with backoff and jitter, because you’re a good citizen. Payments gets sick one afternoon. Each checkout process fails, waits, retries, fails again, waits longer, retries. Polite. Reasonable.
You probably assume that if every caller backs off politely, the sick service gets the breathing room it needs. It’s the opposite. Multiply that polite behaviour by ten thousand checkout processes and payments is being held underwater by a tide of retries from a fleet that was carefully told to be nice about it. Each individual caller is well-behaved; the aggregate is a sustained denial-of-service against a service that’s trying to recover. The slower payments gets, the more requests pile up in checkout’s threadpool or event loop, and the slower checkout gets — which often takes checkout down with it. This failure shape has a name: cascading failure.
Backoff is local. Each call decides independently when to try again. Nothing in the system has the job of noticing that the answer is clearly “down” and stopping the asking on everyone’s behalf. That’s the gap a circuit breaker fills. It sits in front of the call site, watches the recent error rate, and when failure crosses a threshold, it stops making the call at all for a window. Checkout requests in that window fail fast — locally, without touching the network — so payments gets quiet air and checkout doesn’t pile up work it can’t finish.
The metaphor is the household kind: an electrical circuit breaker doesn’t fix a short circuit. It refuses to keep delivering current into one, which is what stops the wiring from catching fire. Same instinct here: when something downstream is wrong, stop pushing into it.
Why it matters now
Anywhere a service has many callers and at least one downstream that can be slow, breakers show up:
- Service meshes. Envoy splits the role across two features.
Circuit breaking
enforces cluster-level limits (max connections, pending requests,
requests, retries). Outlier detection
ejects individual unhealthy hosts. Together they play the breaker role
at the proxy. Istio exposes both via
connectionPoolandoutlierDetection; Linkerd has its own endpoint-level failure-accrual model — same goal, not the same knobs. - RPC frameworks. gRPC’s service config supports either a
retryPolicyor ahedgingPolicyper method (one or the other, not both), and in-house RPC stacks frequently add a breaker layer on top of that. The AWS SDKs don’t ship a generic breaker; their retry story is built on a retry-quota / client-side rate-limiting model instead (the exact defaults differ per SDK and retry mode — check the one you use). The documented AWS story is about capping and pacing retries at the client, not about a generic breaker object you configure. - Browser-side fetch wrappers and mobile clients. When the network is flaky, a breaker on the device prevents UI threads from queuing dozens of doomed retries while the user keeps tapping — the same logic as the server-side version, applied to one device’s own queue of doomed work.
- LLM and agent code. An agent that calls a model provider in a loop has the same shape as checkout calling payments, minus the retry discipline: it can burn through quota retrying 5xx for a provider that’s plainly degraded. A breaker around the provider call gives the loop a way to fail fast and take the alternate path (a smaller model, a cached answer, a graceful “try again later”) instead of stalling.
- Database connection pools. A pool with a bounded size and a short acquisition timeout is doing a cruder version of the same job: it caps how many request threads can be stuck waiting on a sick database at once. Without some such bound, a brief DB blip hangs every request thread and the web tier dies before the DB recovers. (Pool libraries differ in how much breaker-like behaviour they add on top, so the pool’s own docs are the answer here rather than the general pattern.)
The cost of getting it wrong is the cost of the cascading-failure postmortem: you discover that the primary service didn’t really go down — its dependency did, and your service held the door open for the fire to walk through.
The short answer
circuit breaker = error counter + state machine + cooldown window
Picture to keep: a bouncer standing at checkout’s door to payments, who stops letting anyone through the moment the room catches fire — and every so often cracks the door to send one person in to check whether it’s still burning.
A breaker watches the recent results of a call. When errors cross a threshold, it opens — meaning subsequent calls fail immediately without going to the network. After a cooldown, it goes half-open, letting a small probe through. If the probe succeeds, it closes and normal traffic resumes; if it fails, the cooldown starts again. Three states, one counter, one timer. The whole pattern.
How it works
Build it up the way you’d have to build it yourself, one broken version at a time.
Attempt 1: just stop calling
Checkout counts failures against payments. Past some threshold, stop calling — return an error locally instead. That’s the whole idea, and it already solves the original problem: payments stops receiving the flood.
Why it breaks: you’ve now permanently disabled payments. Nothing in this design ever tries again, so a thirty-second blip becomes an outage that ends when a human notices.
Attempt 2: add a cooldown
So set a timer. Stop calling for, say, thirty seconds, then resume normally. Now you have two states — closed (calls pass through, outcomes get counted) and open (calls short-circuit locally, timer counting down).
Why it breaks: when the timer expires you resume all traffic at once, with no evidence payments recovered. If it hasn’t, ten thousand checkout processes hit it simultaneously and you’ve re-created the flood you opened the breaker to stop — a thundering herd aimed at a service that was mid-recovery. You either flap like this or you stay open far too long out of caution.
Attempt 3: probe before committing
Add a third state. When the cooldown elapses, go half-open: let one request through as a probe while everything else keeps failing fast. If the probe succeeds, close and resume normal traffic. If it fails, go back to open and restart the cooldown.
flowchart LR
C[CLOSED] -->|error rate over threshold| O[OPEN]
O -->|cooldown elapses| H[HALF_OPEN]
H -->|probe succeeds| C
H -->|probe fails| O
That’s the pattern. The half-open state is the easiest part to leave out when you build one yourself, and it’s the part that answers is it back? with evidence instead of a timer. Letting one request through and gating the verdict on it is the cheapest experiment that answers the question.
Why it still breaks: “let one request through” needs enforcing. Without
a concurrency cap, every waiter races through the instant the state flips
and you’re back to the herd. Breaker libraries expose this as a
setting — Resilience4j calls it permittedNumberOfCallsInHalfOpenState and
defaults it to ten — and that number is exactly the difference between a
probe and a small stampede, so it is worth setting deliberately rather than
inheriting.
Attempt 4: count the right things
The state machine is now correct and still useless if it’s counting the wrong events. Two calibrations matter:
- Which errors count. A 500 from the downstream counts. A timeout counts. A 404 for a thing the user asked about almost certainly does not — it’s a successful answer to a question. Lumping 4xx in with 5xx is a classic miscalibration: the breaker opens on a hot product page that returns a lot of 404s and takes the whole feature down.
- How “rate” is measured. Two shapes are common: a rolling window (“more than 50% of the last 100 calls failed”) or a consecutive-failure counter (“five failures in a row”). The window form is more robust to bursty workloads; the counter form is simpler and cheaper. Pick one, document the threshold, and make sure there’s a minimum sample size — opening because three out of three calls failed when the service had three calls all minute is how you produce a self-inflicted outage.
Attempt 5: give each destination its own breaker
One breaker for “the outside world” seems tidy until inventory has a bad minute and trips the breaker that checkout uses for payments. A breaker is a per-destination object, not a per-call or per-process one: one breaker per downstream you want to protect, typically per (service, endpoint) or (service, region) tuple. Share one across unrelated dependencies and you get either over-tripping (one bad backend opens the whole world) or under-tripping (payments’ healthy traffic disguises inventory’s failures).
In a service mesh, the breaker lives in the proxy — a
sidecar
next to each instance, or a shared per-node proxy in the sidecarless
deployment modes — keyed by upstream cluster. In application code, it’s
typically a singleton per logical client (paymentsClient.breaker,
inventoryClient.breaker).
Fallbacks: the part the diagram doesn’t show
A breaker that just throws a “circuit open” error is half a feature. The other half is what the caller does instead. Common shapes:
Staying with checkout:
- Cached or stale answer. “Show the last known shipping estimate and mark it as potentially stale.” Right when freshness is nice-to-have.
- Degraded mode. “Skip the recommendation widget, render the rest of the cart.” Right when one feature shouldn’t take down the surface.
- Smaller / cheaper backend. For a call with a cheaper substitute — a large model with a small one behind it, a rich search with a keyword fallback — degrade the quality rather than the availability.
- Surface the failure honestly. “We can’t process payments right now; please try in a minute.” Sometimes the right answer is an error — but a fast, clear one beats a 30-second timeout.
The fallback is application-specific, which is why most generic libraries make you provide it. The breaker handles the when; you handle the what.
Show the seams
- Backoff and breakers solve different problems. Backoff is “this one call should wait before trying again.” A breaker is “the system should stop trying for a window.” A correctly-built client uses both: jittered backoff for the individual call, a breaker for the aggregate of calls against one downstream. The two posts in related below are companions, not alternatives.
- Breakers can mask problems. A breaker that opens on every minor blip will hide real signal from your dashboards if you’re not careful; the dashboard sees “low error rate from this service” and never asks why traffic also dropped. Track open/half-open transitions explicitly, not just downstream error rates.
- Don’t trip on cold starts. A breaker that observes the first three calls of a freshly-deployed pod and opens because the connection pool is still warming up will gate the whole pod’s traffic on a startup artifact. A “minimum throughput in the window” guard fixes this.
- Open isn’t always the safe default. For idempotent reads, failing fast is usually fine. For a write that the user has already paid the latency cost to start — say, a checkout — a circuit-open error during checkout submission may be worse for the business than one extra slow attempt. Tune by call, not by service.
You started with circuit breaker = error counter + state machine + cooldown window. What did checkout’s flood actually force into that line? — + a half-open probe with a concurrency cap. Without it the “cooldown window”
just schedules the next flood, and the state machine has two states instead
of three. The counter tells you when to stop calling; only the probe can
tell you when it’s safe to start again.
Famous related terms
- Exponential backoff —
retry policy = exponential backoff + jitter + stopping rule— the per-call timing rule. Backoff and breakers compose; they don’t replace each other. - Idempotency keys —
idempotency key = client-generated unique ID + server-side dedupe table— what makes the retries that do slip past the breaker safe to reissue. - Bulkhead —
bulkhead = per-dependency concurrency limit + isolated resource pool— a sibling pattern that bounds how much of your fleet any one slow dependency can swallow. Breakers stop calls; bulkheads quarantine them. - Hedged requests —
hedged request = primary call + duplicate sent after a small delay— the opposite instinct, used for tail-latency-sensitive reads. Hedge healthy services; break tripping ones. - Load shedding —
load shedding = drop excess traffic + protect the rest— the server-side cousin. The downstream sheds when overloaded; the caller’s breaker is what notices and stops feeding it. - Outlier detection (Envoy) —
outlier detection = per-host failure tracking + ejection from the load-balancing set for a growing window— the per-instance sibling of cluster-level circuit breaking, with its own knobs (max ejection percent, base ejection time that grows with repeat ejections).
Going deeper
- Michael Nygard, Release It! — go here for the original statement of the pattern, if you want to know what problem its author thought he was solving rather than what the libraries later made of it.
- Martin Fowler’s “CircuitBreaker” article — the explainer to read if you want the state machine walked through slowly, with code, before you pick a library.
- Envoy’s documentation on circuit breaking and outlier detection — the rabbit hole: what knobs a production implementation actually exposes (max connections, max pending, max requests, consecutive 5xx, ejection percentage), which is where the tuning questions above become concrete.