Why does TCP have congestion control?
The internet didn't always have it. Once, in 1986, it nearly fell over. The fix wasn't a protocol change — it was endpoints learning to back off.
On this page
The picture version
Five pictures for a reader who has never thought about how a shared link stays usable, following one evening: hotel Wi-Fi at 9 p.m.
1 · The problem
A hundred guests, one uplink, and the page still loads.
2 · Without it
Retransmitting harder is a loop that feeds itself.
3 · The shape of the fix
Probe upward fast, feel forward slowly, cut hard on loss.
4 · Where the assumption cracks
The whole design reads loss as “the network is full.”
5 · Keep this card
The whole thing on one index card.
Why it exists
Hotel Wi-Fi, 9 p.m. Full signal bars, and a plain web page takes twenty seconds. Nothing is broken: a hundred guests are sharing one uplink out of the building, and it is simply full. Notice what doesn’t happen, though — the page eventually loads. Slow, but it arrives. Every laptop in that building is quietly holding itself back so the shared pipe stays usable. Hold onto that hotel at 9 p.m.; it’s the example the rest of this post runs on.
Now imagine the opposite reflex. Picture a four-lane jam where every driver’s response to slowing down is to honk and accelerate. The jam gets worse, more honking, and soon the road carries zero useful traffic — full of cars, nobody moving. The analogy breaks in one important place: cars can’t be duplicated, but a retransmitted packet is a new copy of an old one, so a congested network doesn’t just slow down — it fills with work it will throw away. The early internet ran into exactly that in 1986. The fix wasn’t to widen the road or add traffic lights. It was to teach every endpoint to voluntarily slow down when it sensed congestion, and speed back up when the road cleared.
In the mid-1980s the early internet briefly stopped working. The standard story — and this one is well-attested — is that the link between LBL and UC Berkeley, a few hundred yards apart, dropped from 32 kbit/s to about 40 bit/s — a factor of roughly a thousand, reported in Jacobson’s 1988 paper. Not because anything broke, but because every sender was doing exactly what TCP told them to: when a packet is lost, retransmit it.
The problem is what happens when many senders all do that at once. A router’s queue fills up, packets get dropped, every sender notices and retransmits, those retransmits also get dropped, and the senders retransmit again. The network is now spending all of its capacity carrying packets that will be thrown away. This is congestion collapse — a stable equilibrium where useful throughput approaches zero while the wires stay completely full.
Van Jacobson’s 1988 paper Congestion Avoidance and Control is the canonical fix. The remarkable thing is what it didn’t change: not the IP layer, not the routers, not the protocol on the wire. Just the sender’s sending logic. Endpoints learned to slow down on their own, by reading loss as a signal that the network was full.
Why it matters now
Every TCP connection you open today — every git push, every database connection, every HTTP/1.1 and HTTP/2 request — runs a congestion control algorithm in the kernel. HTTP/3 doesn’t escape it: QUIC reimplements the same ideas in userspace instead. Datacenter networks, satellite links, and your phone’s LTE connection all depend on senders voluntarily holding back to keep the shared substrate usable.
The choice of algorithm is also a live area: Linux has used CUBIC as its default for many years; Google has publicly documented deploying BBR across its own services, including YouTube and Google Cloud, along with the throughput wins it saw — though not a current production-share breakdown. Datacenter operators, meanwhile, run variants built on ECN marking rather than loss. The mechanism that saved the internet in 1988 is still the surface where most of the interesting transport-layer engineering happens.
The short answer
TCP congestion control = a window that grows when ACKs arrive + shrinks when loss happens
Picture to keep: each laptop in the hotel holding a fistful of letters it’s allowed to have in the mail at once — the fistful grows a little with every reply that comes back, and gets cut in half the moment a letter goes missing.
A sender keeps a congestion window — how many bytes it’s allowed to have in flight at once. Every successful ACK is evidence that the network had room, so the window grows. Every dropped packet is evidence that the network is full, so the window shrinks. The clever part is the shape of grow and shrink, because the network is shared and you want every sender to converge to a fair share without coordinating.
How it works
Derive it the way the constraints force it. Your laptop has just joined the hotel network and wants to load a page.
Naive attempt: send as fast as the link card allows. This is the pre-1988 behavior, and it’s what collapsed. Your laptop has no idea what share of that uplink is available; guessing high means dumping packets into a queue that’s already full, and every one of those drops becomes a retransmission — more load, caused by overload.
Fix: don’t guess, probe. Start with a tiny number of bytes in flight and double it every RTT while ACKs keep coming back. That’s slow start — historically from 1 segment; modern stacks commonly begin at an initial window of 10. The name is a misnomer: doubling is exponential, and it’s the fastest sane way to find an unknown ceiling from below.
But doubling always overshoots. By definition you learn the ceiling by hitting it, and hitting it means a full queue and a lost packet — the very thing you were avoiding. Fix: congestion avoidance. Once the window passes a threshold (roughly the last known good size), switch from doubling to adding one segment per RTT. Near the cliff, feel forward.
Something still has to happen when a packet does go missing. Old TCP’s answer — retransmit — was the collapse mechanism. Fix: treat loss as a message from the network. On a drop, the sender retransmits and cuts its window (classic TCP Reno halves it) and resets the threshold. That single inversion, retransmit-and-slow-down instead of retransmit-harder, is what breaks the collapse loop. Note what it assumes, because the whole edifice rests on it: that a lost packet is decent evidence of a full queue.
But the cheapest way to notice loss is the worst one. A retransmission timeout is deliberately set well above the measured RTT — it has to be, or normal jitter would trigger spurious retransmits — so waiting for one means stalling for far longer than a round trip and then restarting from a tiny window. That’s a twenty-second page load right there. Fix: fast retransmit and fast recovery. Duplicate ACKs (“still missing the same byte”) tell the sender about a drop within roughly one RTT, so it can resend and continue from a halved window instead of from scratch.
The overall shape is sometimes called AIMD: add a little when things are going well, multiply by a fraction when they’re not. It looks arbitrary, but it has a real property: when several AIMD senders with comparable round-trip times share a bottleneck, they converge toward roughly equal shares of it, without ever talking to each other. The link itself becomes the coordination channel — a full queue is the message. The fairness result is genuinely conditional, though; unequal RTTs break it, as the questions at the end of this post get into.
The seam to look at: this whole mechanism treats packet loss as congestion. On the wired networks of 1988 that was a good approximation — most drops really were full queues. It’s a much worse one on a Wi-Fi link or an LTE radio, where packets are routinely lost to interference. A sender on a flaky café connection halves its window in response to noise, not congestion, and throughput tanks. That mismatch is a large part of BBR’s motivation: rather than treating loss as the primary signal, BBR estimates the path’s bottleneck bandwidth and its minimum round-trip time and paces itself to that model. Whether that’s actually better is contested, and it depends on what’s sharing the link with you. Worth stating a boundary here too: there’s no reliable public census of which algorithms carry what share of internet traffic — public sources mostly expose defaults and individual deployment reports, not a measured breakdown.
You started with congestion control = window that grows on ACKs + shrinks on loss. What did the derivation add that the one-liner hides? — + loss is an inference, not a measurement. Nobody tells your laptop the hotel uplink is full; it guesses that from a missing packet, and every argument about CUBIC vs. BBR vs. ECN is an argument about whether that guess is still a good one.
Check yourself
Before you go — your video call is fine until your housemate starts a big upload, and then the call gets choppy even though the router reports no lost packets at all. What’s happening, and why doesn’t loss-based congestion control fix it?
Answer
The upload’s window keeps growing because it never sees a drop — and it never sees a drop because the router has a very large buffer, so excess packets get queued rather than discarded. Loss-based congestion control only backs off on loss, so it happily fills that buffer to the brim. Your call’s packets now wait behind a deep queue: no loss, but big latency. That’s bufferbloat, and the fixes are in the queue (smarter queue management, ECN marking) or in a sender that watches delay rather than loss.
And one more — two flows share a link, one with a 10 ms RTT and one with a 200 ms RTT, both running classic AIMD. Do they end up with equal shares?
Answer
No. Classic AIMD grows the window by about one segment per RTT, so the short-RTT flow probes for more bandwidth far more often and claims a larger share. AIMD’s fairness result is about flows with comparable RTTs sharing a bottleneck; RTT unfairness is a known, well-documented weakness, and it’s one of the things later algorithms were designed to reduce — CUBIC, for instance, grows its window as a function of the time since the last congestion event rather than per-ACK, which lessens the RTT bias without eliminating it.
Famous related terms
- Bufferbloat —
bufferbloat = oversized router buffers + loss-based congestion control— when a router’s queue is huge, TCP keeps filling it because it never sees a drop, and latency through the queue balloons. The reason your video call stutters when someone else starts an upload. - ECN —
ECN = explicit congestion notification— routers mark packets instead of dropping them. A way to tell senders “slow down” without the brutality of a packet loss. Underused on the public internet; widely used inside datacenters. - BBR —
BBR ≈ estimate the path's bandwidth and min-RTT, then pace to that model— a different philosophy from AIMD, where loss stops being the primary signal (though it still drives recovery). Deployed at Google; performance vs. CUBIC depends heavily on the workload, and the literature on fairness between BBR and loss-based flows is not flattering in every scenario. - QUIC —
QUIC = UDP + TCP-style reliability + TLS in userspace— moves congestion control out of the kernel, making it much easier to ship a new algorithm.
Going deeper
- Van Jacobson, Congestion Avoidance and Control (SIGCOMM 1988) — the primary source, and the founding question: what exactly went wrong in 1986, and why was changing only the senders enough to fix it?
- RFC 5681 — the normative answer to “what precisely does slow start / congestion avoidance / fast retransmit / fast recovery do,” in the form an implementer would need.
- The BBR paper plus its follow-up literature — the rabbit hole for whether loss is still the right congestion signal; read the fairness critiques alongside it rather than the summary.