Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do CDNs exist when we already have fast servers?

Your origin server can be the fastest box on Earth and your users in São Paulo will still hate it. CDNs exist because the speed of light, not your CPU, is the bottleneck.

Networking intro Apr 29, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never run a website, following one request: your Sydney colleague loading example.com/logo.png.

1 · The problem

The same server. Wildly different experiences.

one server in a Virginia data center, serving the world round trip to the origin New York feels instant London fine Lagos painful Sydney a couple of seconds per page You upgrade the box. Nothing changes. light through glass over the great-circle distance is already ~150 ms round trip — before any of your code runs
Nothing is broken and the server isn’t loaded. The bill is distance, and no amount of CPU pays it down — which is why the fix has to be geographic rather than computational. Bar lengths here are illustrative, not measured.

2 · The naive fix, escalating

Every extra server fixes exactly one complaint.

the naive fix, and the one after it, and the one after that rent a box in Sydney Lagos still waits rent one in Lagos too São Paulo still waits … and Frankfurt, and … now you operate a global fleet a CDN is what you get when somebody else runs that fleet amortised across thousands of customers, as dozens to hundreds of standing locations — points of presence each one runs cache nodes; the goal is to serve logo.png from the closest that has it
Capacity planning, deploys, certificates, and a decision about which box each visitor should talk to — that is the fleet you just built by accident. A CDN is that fleet as a service, which is the only reason the economics work for a single site.

3 · Finding the near one

She typed one hostname. Something has to steer her.

she typed one hostname — so how does her laptop reach the Sydney POP? GeoDNS the DNS answer differs by who is asking same name in, nearby IP out cheap, broadly compatible but it sees the resolver, not the user Anycast many locations advertise the same IP address the internet’s own routing delivers to the nearest one but you don’t get to pick which one someone lands on Many CDNs combine both. “All CDNs use anycast” is not true.
Both blind spots are real and neither is fixable by trying harder: GeoDNS answers the resolver rather than the person behind it, and anycast lands you wherever routing decided this hour. Steering is a best guess about proximity, not a measurement of it.

4 · The dependency everything rests on

A nearby server with nothing on it is just an extra hop.

steering only pays off if the POP can answer HIT returned immediately; the origin is never contacted MISS fetch from origin (or a regional parent cache), store it, return it the next user in that region gets a hit so for the caching half of the story, the hit ratio is the whole game 50% hits misses every miss pays the full origin round trip plus the user-to-POP hop A CDN in front of uncacheable responses buys far less than the brochure suggests.
The other failure is the opposite one: publish something that goes viral and every POP misses at the same instant, so the origin takes one request per location in the same second. The answer is a hierarchy — edge POPs in front of regional shields in front of origin — so the misses fan in rather than fanning out.

5 · Keep this card

The whole thing on one index card.

CDN = a global fleet of caches + smart routing to the nearest one ∴ the win is latency, paid for in cache hits the CPU was never the bottleneck — geography was, and cacheability is the price of fixing it
Picture to keep: not one warehouse shipping worldwide, but a corner shop in every city stocked from that warehouse — and a doorman who sends each customer to their nearest shop. Where it breaks: the shops don’t restock on a schedule, they only fetch an item the first time someone in that city asks for it.

Why it exists

You ship a site, it feels instant on your laptop, and then a colleague in Sydney says every page takes a couple of seconds. Nothing is broken. The server isn’t loaded. You upgrade the box anyway and it changes nothing — because the thing you’re fighting isn’t the server.

Picture the setup: one server in a Virginia data center, serving the world. Page loads great if you’re in New York. It’s fine in London. It’s noticeably slow in Sydney. It’s painful in Lagos. That Sydney colleague loading example.com/logo.png is the example this post follows.

You can’t fix this by buying a faster server. The Sydney user’s laptop and your Virginia server are on opposite sides of the planet, and a packet has to physically traverse fiber across an ocean. The physics floor — light through glass over the great-circle distance — is on the order of 150 ms round trip; real-world routes through real cables typically come in higher than that, before any of your code runs. Modern protocols still spend round trips setting up a connection (TCP plus TLS), and a typical page then fires off dozens of asset requests. Connection reuse and multiplexing take some of that back, but the distance is charged on every exchange that has to reach the origin. Latency stacks.

A CDN — a Content Delivery Network — solves a problem your origin physically cannot: it puts a copy of your stuff near the user, so the round trip is short. In that scenario the CPU was never the bottleneck. Geography was.

Why it matters now

A large fraction of the traffic you experience as “the web” is fronted by a CDN, even when it doesn’t look like one — static assets, video, software updates, API responses, edge-rendered pages, package registry tarballs, container image layers, model weights. The exact share depends on what you measure (HTML requests, third-party requests, total bytes); the HTTP Archive Web Almanac publishes yearly numbers and they vary considerably by category, but the direction is consistent: heavy CDN involvement, especially for third-party and asset traffic.

For engineers in the AI era specifically:

The short answer

CDN = global fleet of caches + smart routing to the nearest one

Picture to keep: not one warehouse shipping worldwide, but a corner shop in every city stocked from that warehouse — and a doorman who sends each customer to their nearest shop. Where the picture breaks: the shops don’t restock on a schedule, they only fetch an item the first time someone in that city asks for it.

A CDN is many servers, in many cities, each holding a copy of your content, with a system out front that steers each user to a nearby copy. Your Sydney colleague talks to a machine 20 ms away instead of 200 ms away, and your origin only sees a trickle of cache misses.

How it works

Naive fix: rent a second server in Sydney. This genuinely works — it fixes the exact complaint you got. It just fixes only that one — your colleague in Lagos still waits, so you rent one there too, and in São Paulo, and in Frankfurt, and now you’re operating a global fleet: capacity planning, deploys, certificates, and a decision about which box each visitor should talk to. A CDN is what you get when someone else runs that fleet and amortizes it across thousands of customers.

So one extra server doesn’t scale; you need standing ones everywhere. Those locations are called points of presence (POPs). A CDN operates servers in dozens to hundreds of locations worldwide — major exchange points, ISP facilities, metro data centers. Each POP runs cache nodes. When your Sydney colleague requests https://example.com/logo.png, the goal is to serve logo.png from the closest POP that has it.

But how does her laptop find the Sydney POP? She typed one hostname, and example.com has to resolve to something — the naive answer, one IP for one machine, is the whole problem you’re trying to escape. Two mechanisms dominate, often combined:

Many CDNs combine these — for example, anycast at the network edge plus DNS-based steering, and internal routing to push the request to a POP that’s warm and healthy. The exact mix varies by provider; “all CDNs use anycast” is not true.

But a nearby server with nothing on it is just an extra hop. Steering only pays off if the POP can answer. So it caches — and the moment it does, you inherit every hard question about when a copy is still valid. When the chosen POP gets the request, it checks its cache:

Once a POP holds a copy, the question stops being “can this be cached?” and becomes “when may that copy still be reused?” What’s cacheable, and for how long, is controlled by HTTP cache headers plus CDN-specific config. Three different concerns, often muddled together:

Static files are easy. Dynamic responses get harder, which is where features like stale-while-revalidate, surrogate keys for targeted purges, and edge-computed personalization come in.

And a cold cache everywhere at once is its own failure. Publish something that goes viral and every POP misses simultaneously, so your origin takes one request per POP — or worse, per node — in the same second. Fix: a cache hierarchy. Edge POPs sit in front of regional shields, which sit in front of origin, so the misses fan in rather than fanning out.

Show the seams

A few things the simple story glosses over:

You started with CDN = global fleet of caches + smart routing to the nearest one. What did your Sydney colleague add? — + the classic speedup is for latency, and it works best to the extent you get cache hits. That single dependency explains the rest: why hit ratio is the metric operators watch, why hashed filenames became standard practice, why origin shields exist, and why a CDN in front of uncacheable, personalized responses buys you far less than the brochure suggests.

Going deeper