Why do CDNs exist when we already have fast servers?
Your origin server can be the fastest box on Earth and your users in São Paulo will still hate it. CDNs exist because the speed of light, not your CPU, is the bottleneck.
On this page
The picture version
Five pictures for a reader who has never run a website, following one request: your Sydney colleague loading example.com/logo.png.
1 · The problem
The same server. Wildly different experiences.
2 · The naive fix, escalating
Every extra server fixes exactly one complaint.
3 · Finding the near one
She typed one hostname. Something has to steer her.
4 · The dependency everything rests on
A nearby server with nothing on it is just an extra hop.
5 · Keep this card
The whole thing on one index card.
Why it exists
You ship a site, it feels instant on your laptop, and then a colleague in Sydney says every page takes a couple of seconds. Nothing is broken. The server isn’t loaded. You upgrade the box anyway and it changes nothing — because the thing you’re fighting isn’t the server.
Picture the setup: one server in a Virginia data center, serving the world.
Page loads great if you’re in New York. It’s fine in London. It’s noticeably
slow in Sydney. It’s painful in Lagos. That Sydney colleague loading
example.com/logo.png is the example this post follows.
You can’t fix this by buying a faster server. The Sydney user’s laptop and your Virginia server are on opposite sides of the planet, and a packet has to physically traverse fiber across an ocean. The physics floor — light through glass over the great-circle distance — is on the order of 150 ms round trip; real-world routes through real cables typically come in higher than that, before any of your code runs. Modern protocols still spend round trips setting up a connection (TCP plus TLS), and a typical page then fires off dozens of asset requests. Connection reuse and multiplexing take some of that back, but the distance is charged on every exchange that has to reach the origin. Latency stacks.
A CDN — a Content Delivery Network — solves a problem your origin physically cannot: it puts a copy of your stuff near the user, so the round trip is short. In that scenario the CPU was never the bottleneck. Geography was.
Why it matters now
A large fraction of the traffic you experience as “the web” is fronted by a CDN, even when it doesn’t look like one — static assets, video, software updates, API responses, edge-rendered pages, package registry tarballs, container image layers, model weights. The exact share depends on what you measure (HTML requests, third-party requests, total bytes); the HTTP Archive Web Almanac publishes yearly numbers and they vary considerably by category, but the direction is consistent: heavy CDN involvement, especially for third-party and asset traffic.
For engineers in the AI era specifically:
- LLM API providers run their inference fleets across many regions for the same latency reason. The plumbing that gets your request to the nearest region is CDN-shaped.
- Model weights and container images are huge. Pulling a multi-gigabyte image from a single origin on every deploy can saturate the link. Major registries (Docker Hub, GHCR, Hugging Face) front their downloads with CDN-style caching for that reason.
- “Edge functions” — running code on the CDN’s POPs instead of your origin — are the natural extension once you’ve already got a global fleet of servers.
The short answer
CDN = global fleet of caches + smart routing to the nearest one
Picture to keep: not one warehouse shipping worldwide, but a corner shop in every city stocked from that warehouse — and a doorman who sends each customer to their nearest shop. Where the picture breaks: the shops don’t restock on a schedule, they only fetch an item the first time someone in that city asks for it.
A CDN is many servers, in many cities, each holding a copy of your content, with a system out front that steers each user to a nearby copy. Your Sydney colleague talks to a machine 20 ms away instead of 200 ms away, and your origin only sees a trickle of cache misses.
How it works
Naive fix: rent a second server in Sydney. This genuinely works — it fixes the exact complaint you got. It just fixes only that one — your colleague in Lagos still waits, so you rent one there too, and in São Paulo, and in Frankfurt, and now you’re operating a global fleet: capacity planning, deploys, certificates, and a decision about which box each visitor should talk to. A CDN is what you get when someone else runs that fleet and amortizes it across thousands of customers.
So one extra server doesn’t scale; you need standing ones everywhere.
Those locations are called
points of presence
(POPs). A CDN operates servers in dozens to hundreds of locations
worldwide — major exchange points, ISP facilities,
metro data centers. Each POP runs cache nodes. When your Sydney colleague
requests https://example.com/logo.png, the goal is to serve logo.png from
the closest POP that has it.
But how does her laptop find the Sydney POP? She typed one hostname, and
example.com has to resolve to something — the naive answer, one IP for
one machine, is the whole problem you’re trying to escape. Two mechanisms
dominate, often combined:
- GeoDNS.
When the user’s resolver looks up
example.com, the CDN’s DNS server hands back an IP for a nearby POP. Cheap and broadly compatible; the downside is it routes based on the user’s resolver, which isn’t always close to the user. - Anycast. Multiple POPs advertise the same IP address into the global routing table. The internet’s own routing (BGP) delivers the packet to whichever POP is “closest” by routing metric. Latency-aware in a structural way, but you don’t get to pick which POP someone lands on.
Many CDNs combine these — for example, anycast at the network edge plus DNS-based steering, and internal routing to push the request to a POP that’s warm and healthy. The exact mix varies by provider; “all CDNs use anycast” is not true.
But a nearby server with nothing on it is just an extra hop. Steering only pays off if the POP can answer. So it caches — and the moment it does, you inherit every hard question about when a copy is still valid. When the chosen POP gets the request, it checks its cache:
- Hit — return the cached response immediately. Origin is not contacted.
- Miss — fetch from origin (or a regional parent cache), store the response, and return it. The next user in that region gets a hit.
Once a POP holds a copy, the question stops being “can this be cached?” and becomes “when may that copy still be reused?” What’s cacheable, and for how long, is controlled by HTTP cache headers plus CDN-specific config. Three different concerns, often muddled together:
- Freshness —
Cache-Control(and the olderExpires) say how long a response can be served from cache before re-checking with origin. - Validation —
ETagandLast-Modifiedare the tokens used to ask origin “is what I have still good?” without re-downloading the body. - Cache-key variation —
Varytells the cache that the right response depends on a request header (e.g.Accept-Encoding), so it must keep separate cached entries.
Static files are easy. Dynamic responses get harder, which is where features like stale-while-revalidate, surrogate keys for targeted purges, and edge-computed personalization come in.
And a cold cache everywhere at once is its own failure. Publish something that goes viral and every POP misses simultaneously, so your origin takes one request per POP — or worse, per node — in the same second. Fix: a cache hierarchy. Edge POPs sit in front of regional shields, which sit in front of origin, so the misses fan in rather than fanning out.
Show the seams
A few things the simple story glosses over:
- CDNs aren’t only about static files. Modern CDNs terminate TLS, run WAFs, do DDoS scrubbing, route APIs, and execute code at the edge. The cache is the historical core, not the whole product.
- Cache invalidation is the hard part. “Push a new version of
app.js” sounds trivial; doing it consistently across hundreds of POPs in seconds is an actual distributed-systems problem. A common workaround for static assets is to put a content hash in the filename (app.abc123.js) and treat the URL as immutable — change the URL instead of invalidating the cache. - GeoDNS sees the resolver, not the user. If a user in Kenya is using a public resolver hosted in Europe, GeoDNS will route them to a European POP. EDNS Client Subnet (RFC 7871) helps by passing a portion of the client’s subnet to the authoritative server, but it isn’t universally deployed and has its own privacy tradeoffs that not every resolver wants to make.
- Anycast is a routing artifact, not a guarantee. A user can land on a POP that is geographically far but topologically near (or just whatever BGP decided this hour). It usually works; when it doesn’t, debugging is hard because the routing isn’t yours.
- For the caching half of the story, hit ratio is the whole game. A CDN with a 50% hit ratio is doing half that job — every miss pays the full origin round trip plus the user-to-POP hop. The hit ratio depends on how cacheable your content is and how well the cache key is tuned. Typical ratios vary wildly by workload, which is why no single industry number is worth quoting; it’s still the metric operators actually watch.
- CDNs are a centralization story. A small number of providers handle the front-door traffic for a sizeable share of popular websites. Published measurements vary by methodology, so the percentage is worth treating carefully, but the practical evidence is clear: when one big CDN has a bad config push, a noticeable fraction of the internet appears to break at once. That’s not a bug in the CDN model; it’s a structural consequence of it.
You started with CDN = global fleet of caches + smart routing to the nearest one. What did your Sydney colleague add? — + the classic speedup is for latency, and it works best to the extent you get cache hits. That single dependency explains
the rest: why hit ratio is the metric operators watch, why hashed filenames
became standard practice, why origin shields exist, and why a CDN in front of
uncacheable, personalized responses buys you far less than the brochure
suggests.
Famous related terms
- POP (Point of Presence) —
POP = a CDN's local server cluster + the network links into it— the unit a CDN is built out of. - Anycast —
anycast ≈ "same IP advertised from many places, the network picks one"— how packets find a nearby POP without DNS being involved. - GeoDNS —
GeoDNS = DNS + per-resolver answers— older but still common steering mechanism. - Edge function —
edge function ≈ "your code, but running in the POP"— what you get once the CDN already has compute everywhere. - Origin shield —
origin shield = an extra cache tier in front of your origin— absorbs cache misses so origin doesn’t get hammered on cold cache.
Going deeper
- RFC 9111 (HTTP Caching) — the primary source for “what exactly is a cache allowed to do with my response,” which is the question behind most surprising CDN behavior.
- Karger et al., Consistent Hashing and Random Trees (STOC 1997) — read this for the origin question: how do you spread cached content across a changing set of machines without reshuffling everything? One influential building block behind distributed caches, published shortly before Akamai was founded in 1998.
- The HTTP Archive Web Almanac CDN chapter — the rabbit hole for “what is actually deployed out there,” measured across millions of sites rather than asserted.