Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does TCP have congestion control?

The internet didn't always have it. Once, in 1986, it nearly fell over. The fix wasn't a protocol change — it was endpoints learning to back off.

Networking intermediate Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never thought about how a shared link stays usable, following one evening: hotel Wi-Fi at 9 p.m.

1 · The problem

A hundred guests, one uplink, and the page still loads.

~100 guests, 9 p.m. full, all evening one uplink out of the building the internet plenty of capacity Notice what doesn’t happen: the page still loads. every laptop in the building is quietly holding itself back so the shared pipe stays usable
Nothing here is broken and nothing is coordinating. The restraint is voluntary, per-laptop, and invisible — and it is the only reason a full link degrades into a slow page instead of no page.

2 · Without it

Retransmitting harder is a loop that feeds itself.

a router’s queue fills packets get dropped every sender retransmits more load than before and again worse each round congestion collapse the wires stay completely full and useful throughput approaches zero In 1986 one link fell from 32 kbit/s to about 40 bit/s. a factor of roughly a thousand — and nothing was broken
A retransmitted packet is a fresh copy of an old one, so a congested network doesn’t merely slow down — it fills with work it is going to throw away. The fix changed no router and no packet format, only what a sender does when it notices a loss.

3 · The shape of the fix

Probe upward fast, feel forward slowly, cut hard on loss.

bytes in flight (the congestion window) time slow start: double every RTT then add one segment per RTT halve on loss Add a little when things go well; multiply down when they don’t. several senders doing this on one link converge toward roughly equal shares — if their round-trip times are comparable
Doubling is the fastest sane way to find an unknown ceiling from below, and it always overshoots — you learn where the ceiling is by hitting it. The sawtooth is not a defect; it is the shape of a sender that never gets told the answer and has to keep asking.

4 · Where the assumption cracks

The whole design reads loss as “the network is full.”

wired, 1988 most drops really were a full queue somewhere so loss is a good proxy Wi-Fi or mobile packets are routinely lost to interference so the window halves for noise Nobody tells your laptop the uplink is full. It infers that from a missing packet. BBR’s answer: estimate the path’s bottleneck bandwidth and minimum round-trip time, and pace to that model instead whether that is actually better is contested, and it depends on what is sharing the link with you
A sender on a flaky café connection backs off in response to radio noise, not to a full queue, and throughput tanks. The inference was a good approximation on the network it was designed for, and every modern argument about congestion control is an argument about whether it still is.

5 · Keep this card

The whole thing on one index card.

congestion control = a window that grows when ACKs arrive + shrinks when loss happens ∴ loss is an inference, not a measurement every argument about CUBIC vs BBR vs ECN is an argument about whether that inference still holds
Picture to keep: each laptop in the hotel holding a fistful of letters it’s allowed to have in the mail at once — the fistful grows a little with every reply that comes back, and gets cut in half the moment a letter goes missing. Nobody in the building is talking to anybody else; the full link is the only message they share.

Why it exists

Hotel Wi-Fi, 9 p.m. Full signal bars, and a plain web page takes twenty seconds. Nothing is broken: a hundred guests are sharing one uplink out of the building, and it is simply full. Notice what doesn’t happen, though — the page eventually loads. Slow, but it arrives. Every laptop in that building is quietly holding itself back so the shared pipe stays usable. Hold onto that hotel at 9 p.m.; it’s the example the rest of this post runs on.

Now imagine the opposite reflex. Picture a four-lane jam where every driver’s response to slowing down is to honk and accelerate. The jam gets worse, more honking, and soon the road carries zero useful traffic — full of cars, nobody moving. The analogy breaks in one important place: cars can’t be duplicated, but a retransmitted packet is a new copy of an old one, so a congested network doesn’t just slow down — it fills with work it will throw away. The early internet ran into exactly that in 1986. The fix wasn’t to widen the road or add traffic lights. It was to teach every endpoint to voluntarily slow down when it sensed congestion, and speed back up when the road cleared.

In the mid-1980s the early internet briefly stopped working. The standard story — and this one is well-attested — is that the link between LBL and UC Berkeley, a few hundred yards apart, dropped from 32 kbit/s to about 40 bit/s — a factor of roughly a thousand, reported in Jacobson’s 1988 paper. Not because anything broke, but because every sender was doing exactly what TCP told them to: when a packet is lost, retransmit it.

The problem is what happens when many senders all do that at once. A router’s queue fills up, packets get dropped, every sender notices and retransmits, those retransmits also get dropped, and the senders retransmit again. The network is now spending all of its capacity carrying packets that will be thrown away. This is congestion collapse — a stable equilibrium where useful throughput approaches zero while the wires stay completely full.

Van Jacobson’s 1988 paper Congestion Avoidance and Control is the canonical fix. The remarkable thing is what it didn’t change: not the IP layer, not the routers, not the protocol on the wire. Just the sender’s sending logic. Endpoints learned to slow down on their own, by reading loss as a signal that the network was full.

Why it matters now

Every TCP connection you open today — every git push, every database connection, every HTTP/1.1 and HTTP/2 request — runs a congestion control algorithm in the kernel. HTTP/3 doesn’t escape it: QUIC reimplements the same ideas in userspace instead. Datacenter networks, satellite links, and your phone’s LTE connection all depend on senders voluntarily holding back to keep the shared substrate usable.

The choice of algorithm is also a live area: Linux has used CUBIC as its default for many years; Google has publicly documented deploying BBR across its own services, including YouTube and Google Cloud, along with the throughput wins it saw — though not a current production-share breakdown. Datacenter operators, meanwhile, run variants built on ECN marking rather than loss. The mechanism that saved the internet in 1988 is still the surface where most of the interesting transport-layer engineering happens.

The short answer

TCP congestion control = a window that grows when ACKs arrive + shrinks when loss happens

Picture to keep: each laptop in the hotel holding a fistful of letters it’s allowed to have in the mail at once — the fistful grows a little with every reply that comes back, and gets cut in half the moment a letter goes missing.

A sender keeps a congestion window — how many bytes it’s allowed to have in flight at once. Every successful ACK is evidence that the network had room, so the window grows. Every dropped packet is evidence that the network is full, so the window shrinks. The clever part is the shape of grow and shrink, because the network is shared and you want every sender to converge to a fair share without coordinating.

How it works

Derive it the way the constraints force it. Your laptop has just joined the hotel network and wants to load a page.

Naive attempt: send as fast as the link card allows. This is the pre-1988 behavior, and it’s what collapsed. Your laptop has no idea what share of that uplink is available; guessing high means dumping packets into a queue that’s already full, and every one of those drops becomes a retransmission — more load, caused by overload.

Fix: don’t guess, probe. Start with a tiny number of bytes in flight and double it every RTT while ACKs keep coming back. That’s slow start — historically from 1 segment; modern stacks commonly begin at an initial window of 10. The name is a misnomer: doubling is exponential, and it’s the fastest sane way to find an unknown ceiling from below.

But doubling always overshoots. By definition you learn the ceiling by hitting it, and hitting it means a full queue and a lost packet — the very thing you were avoiding. Fix: congestion avoidance. Once the window passes a threshold (roughly the last known good size), switch from doubling to adding one segment per RTT. Near the cliff, feel forward.

Something still has to happen when a packet does go missing. Old TCP’s answer — retransmit — was the collapse mechanism. Fix: treat loss as a message from the network. On a drop, the sender retransmits and cuts its window (classic TCP Reno halves it) and resets the threshold. That single inversion, retransmit-and-slow-down instead of retransmit-harder, is what breaks the collapse loop. Note what it assumes, because the whole edifice rests on it: that a lost packet is decent evidence of a full queue.

But the cheapest way to notice loss is the worst one. A retransmission timeout is deliberately set well above the measured RTT — it has to be, or normal jitter would trigger spurious retransmits — so waiting for one means stalling for far longer than a round trip and then restarting from a tiny window. That’s a twenty-second page load right there. Fix: fast retransmit and fast recovery. Duplicate ACKs (“still missing the same byte”) tell the sender about a drop within roughly one RTT, so it can resend and continue from a halved window instead of from scratch.

The overall shape is sometimes called AIMD: add a little when things are going well, multiply by a fraction when they’re not. It looks arbitrary, but it has a real property: when several AIMD senders with comparable round-trip times share a bottleneck, they converge toward roughly equal shares of it, without ever talking to each other. The link itself becomes the coordination channel — a full queue is the message. The fairness result is genuinely conditional, though; unequal RTTs break it, as the questions at the end of this post get into.

The seam to look at: this whole mechanism treats packet loss as congestion. On the wired networks of 1988 that was a good approximation — most drops really were full queues. It’s a much worse one on a Wi-Fi link or an LTE radio, where packets are routinely lost to interference. A sender on a flaky café connection halves its window in response to noise, not congestion, and throughput tanks. That mismatch is a large part of BBR’s motivation: rather than treating loss as the primary signal, BBR estimates the path’s bottleneck bandwidth and its minimum round-trip time and paces itself to that model. Whether that’s actually better is contested, and it depends on what’s sharing the link with you. Worth stating a boundary here too: there’s no reliable public census of which algorithms carry what share of internet traffic — public sources mostly expose defaults and individual deployment reports, not a measured breakdown.

You started with congestion control = window that grows on ACKs + shrinks on loss. What did the derivation add that the one-liner hides? — + loss is an inference, not a measurement. Nobody tells your laptop the hotel uplink is full; it guesses that from a missing packet, and every argument about CUBIC vs. BBR vs. ECN is an argument about whether that guess is still a good one.

Check yourself

Before you go — your video call is fine until your housemate starts a big upload, and then the call gets choppy even though the router reports no lost packets at all. What’s happening, and why doesn’t loss-based congestion control fix it?

Answer

The upload’s window keeps growing because it never sees a drop — and it never sees a drop because the router has a very large buffer, so excess packets get queued rather than discarded. Loss-based congestion control only backs off on loss, so it happily fills that buffer to the brim. Your call’s packets now wait behind a deep queue: no loss, but big latency. That’s bufferbloat, and the fixes are in the queue (smarter queue management, ECN marking) or in a sender that watches delay rather than loss.

And one more — two flows share a link, one with a 10 ms RTT and one with a 200 ms RTT, both running classic AIMD. Do they end up with equal shares?

Answer

No. Classic AIMD grows the window by about one segment per RTT, so the short-RTT flow probes for more bandwidth far more often and claims a larger share. AIMD’s fairness result is about flows with comparable RTTs sharing a bottleneck; RTT unfairness is a known, well-documented weakness, and it’s one of the things later algorithms were designed to reduce — CUBIC, for instance, grows its window as a function of the time since the last congestion event rather than per-ACK, which lessens the RTT bias without eliminating it.

Going deeper