Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why circuit breakers exist

Backoff makes a single retry polite. But when a downstream is plainly down, every caller in your fleet generously retrying it is the actual problem. A circuit breaker is the small piece that says: stop calling for a while — the answer isn't going to change in the next 50 ms.

Systems intro Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Four pictures for a reader who has already added retries and still went down. The prose below fills in the seams the pictures skip.

1 · The problem

Ten thousand polite callers are still a flood.

ten thousand checkout processes every one backing off politely each one reasonable in isolation payments, trying to recover held underwater and checkout’s threadpool fills with work it can’t finish Backoff is local. Nobody’s job is to notice the answer is clearly “down”. and the slower payments gets, the slower checkout gets — which is how one sick service takes its caller with it
That shape has a name: cascading failure. Each caller decides independently when to try again, so politeness scales the delay but not the decision — and the aggregate is a sustained denial of service against something already struggling.

2 · The move

Put something in front of the call that can stop asking.

your call site the breaker watches the recent error rate healthy → the call goes out tripped → fail here, instantly, no network Two things happen at once, and both matter. payments gets quiet air to recover in — and checkout stops accumulating work it was never going to complete failing fast is not giving up. it is refusing to hold a thread open for an answer you can already predict.
The household metaphor is exact in the part that counts: an electrical breaker doesn’t fix a short circuit, it refuses to keep delivering current into one, which is what stops the wiring catching fire.

3 · The three states

And the middle one is the whole design.

closed calls pass through, failures get counted too many open every call fails instantly, for a set window time’s up half-open let a few calls through and watch what happens it worked → close, and resume normally it failed → open again Without half-open, you either stay shut forever or stampede back all at once.
Those trial calls are the cheapest possible probe: they ask the only question that matters using a handful of requests instead of ten thousand. Send the whole fleet back the instant the timer expires and you have rebuilt the flood you just escaped.

4 · Keep this card

The whole thing on one index card.

circuit breaker = error counter + state machine + cooldown window count what the last few calls did + fail fast once too many of them failed + after a wait, one trial call decides whether to resume ∴ backoff paces one caller. this one decides for all of them.
Picture to keep: the household breaker. It doesn’t repair the short — it refuses to keep feeding current into one, and that refusal is what saves the wiring. Your breaker doesn’t fix payments either. It just stops your fleet from being the reason payments can’t get up.

Why it exists

Your checkout service calls a payments service on every order. You’ve already shipped retries with backoff and jitter, because you’re a good citizen. Payments gets sick one afternoon. Each checkout process fails, waits, retries, fails again, waits longer, retries. Polite. Reasonable.

You probably assume that if every caller backs off politely, the sick service gets the breathing room it needs. It’s the opposite. Multiply that polite behaviour by ten thousand checkout processes and payments is being held underwater by a tide of retries from a fleet that was carefully told to be nice about it. Each individual caller is well-behaved; the aggregate is a sustained denial-of-service against a service that’s trying to recover. The slower payments gets, the more requests pile up in checkout’s threadpool or event loop, and the slower checkout gets — which often takes checkout down with it. This failure shape has a name: cascading failure.

Backoff is local. Each call decides independently when to try again. Nothing in the system has the job of noticing that the answer is clearly “down” and stopping the asking on everyone’s behalf. That’s the gap a circuit breaker fills. It sits in front of the call site, watches the recent error rate, and when failure crosses a threshold, it stops making the call at all for a window. Checkout requests in that window fail fast — locally, without touching the network — so payments gets quiet air and checkout doesn’t pile up work it can’t finish.

The metaphor is the household kind: an electrical circuit breaker doesn’t fix a short circuit. It refuses to keep delivering current into one, which is what stops the wiring from catching fire. Same instinct here: when something downstream is wrong, stop pushing into it.

Why it matters now

Anywhere a service has many callers and at least one downstream that can be slow, breakers show up:

The cost of getting it wrong is the cost of the cascading-failure postmortem: you discover that the primary service didn’t really go down — its dependency did, and your service held the door open for the fire to walk through.

The short answer

circuit breaker = error counter + state machine + cooldown window

Picture to keep: a bouncer standing at checkout’s door to payments, who stops letting anyone through the moment the room catches fire — and every so often cracks the door to send one person in to check whether it’s still burning.

A breaker watches the recent results of a call. When errors cross a threshold, it opens — meaning subsequent calls fail immediately without going to the network. After a cooldown, it goes half-open, letting a small probe through. If the probe succeeds, it closes and normal traffic resumes; if it fails, the cooldown starts again. Three states, one counter, one timer. The whole pattern.

How it works

Build it up the way you’d have to build it yourself, one broken version at a time.

Attempt 1: just stop calling

Checkout counts failures against payments. Past some threshold, stop calling — return an error locally instead. That’s the whole idea, and it already solves the original problem: payments stops receiving the flood.

Why it breaks: you’ve now permanently disabled payments. Nothing in this design ever tries again, so a thirty-second blip becomes an outage that ends when a human notices.

Attempt 2: add a cooldown

So set a timer. Stop calling for, say, thirty seconds, then resume normally. Now you have two states — closed (calls pass through, outcomes get counted) and open (calls short-circuit locally, timer counting down).

Why it breaks: when the timer expires you resume all traffic at once, with no evidence payments recovered. If it hasn’t, ten thousand checkout processes hit it simultaneously and you’ve re-created the flood you opened the breaker to stop — a thundering herd aimed at a service that was mid-recovery. You either flap like this or you stay open far too long out of caution.

Attempt 3: probe before committing

Add a third state. When the cooldown elapses, go half-open: let one request through as a probe while everything else keeps failing fast. If the probe succeeds, close and resume normal traffic. If it fails, go back to open and restart the cooldown.

flowchart LR
    C[CLOSED] -->|error rate over threshold| O[OPEN]
    O -->|cooldown elapses| H[HALF_OPEN]
    H -->|probe succeeds| C
    H -->|probe fails| O

That’s the pattern. The half-open state is the easiest part to leave out when you build one yourself, and it’s the part that answers is it back? with evidence instead of a timer. Letting one request through and gating the verdict on it is the cheapest experiment that answers the question.

Why it still breaks: “let one request through” needs enforcing. Without a concurrency cap, every waiter races through the instant the state flips and you’re back to the herd. Breaker libraries expose this as a setting — Resilience4j calls it permittedNumberOfCallsInHalfOpenState and defaults it to ten — and that number is exactly the difference between a probe and a small stampede, so it is worth setting deliberately rather than inheriting.

Attempt 4: count the right things

The state machine is now correct and still useless if it’s counting the wrong events. Two calibrations matter:

Attempt 5: give each destination its own breaker

One breaker for “the outside world” seems tidy until inventory has a bad minute and trips the breaker that checkout uses for payments. A breaker is a per-destination object, not a per-call or per-process one: one breaker per downstream you want to protect, typically per (service, endpoint) or (service, region) tuple. Share one across unrelated dependencies and you get either over-tripping (one bad backend opens the whole world) or under-tripping (payments’ healthy traffic disguises inventory’s failures).

In a service mesh, the breaker lives in the proxy — a sidecar next to each instance, or a shared per-node proxy in the sidecarless deployment modes — keyed by upstream cluster. In application code, it’s typically a singleton per logical client (paymentsClient.breaker, inventoryClient.breaker).

Fallbacks: the part the diagram doesn’t show

A breaker that just throws a “circuit open” error is half a feature. The other half is what the caller does instead. Common shapes:

Staying with checkout:

The fallback is application-specific, which is why most generic libraries make you provide it. The breaker handles the when; you handle the what.

Show the seams

You started with circuit breaker = error counter + state machine + cooldown window. What did checkout’s flood actually force into that line? — + a half-open probe with a concurrency cap. Without it the “cooldown window” just schedules the next flood, and the state machine has two states instead of three. The counter tells you when to stop calling; only the probe can tell you when it’s safe to start again.

Going deeper