Why containers won over VMs
Both promise isolated, reproducible environments. One boots in milliseconds and ships in megabytes; the other boots in seconds and ships in gigabytes. The reason isn't 'containers are lighter VMs' — they're a different kind of thing entirely.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: give it its own machine
- Fix: drop the second kernel — but then the blindfold has to come from somewhere
- But seeing nothing isn’t the same as taking nothing
- But shipping a root filesystem per service is VM-sized again
- And that’s why there’s nothing left to wait for
- Show the seams
- Check yourself
- Famous related terms
- Going deeper
The picture version
The whole idea in six pictures, for a reader who has typed docker run but never asked what it actually did. The prose below fills in the seams.
1 · The problem
The same request, sixteen years apart.
2 · The old way
A VM duplicates the whole machine, kernel and all.
3 · The trick
A blindfold and a meter, issued to one ordinary process.
4 · The other half of the win
The image is a stack, so you only ship the top of it.
5 · The seam
The blindfold is issued by the host’s own kernel.
6 · Keep this card
The whole thing on one index card.
Why it exists
You need a Postgres for a feature branch. You type docker run postgres, and
before you’ve finished alt-tabbing back to your editor it’s accepting
connections. When you’re done you delete it and the machine looks like it
never happened. That thirty-second Postgres is the running example for this
post — keep it in mind, because the same request in 2010 meant downloading a
multi-gigabyte disk image, booting a whole operating system inside your
computer, and waiting.
That’s what deploying a service used to mean: a virtual machine — a full guest operating system, kernel and all, running on top of a hypervisor on top of the host kernel. It worked, but it was heavy. A “small” service shipped as a multi-gigabyte disk image, took tens of seconds to boot, and spent most of its RAM on a kernel and userland its actual workload would never use. Running ten copies of your service for testing meant ten kernels.
The pain point was a mismatch. What developers actually wanted was: “give me my code, my dependencies, and a filesystem that looks the way I expect — and please don’t let me see anyone else’s stuff.” They didn’t want a second kernel. They wanted isolation of the things above the kernel, not duplication of the kernel.
Linux had been quietly accumulating the pieces to do exactly that:
namespaces
(starting with mount namespaces in 2002; PID, network and the rest arrived
over the following decade, and user namespaces landed in 3.8 / 2013 — later
hardening rather than something the first containers were built on) and
cgroups
(started at Google and shipped in Linux 2.6.24, January 2008). In 2013, Docker packaged those
primitives behind a friendly CLI and an image format you could push and
pull, and over the next several years much of the industry’s
deployment story shifted from VMs to containers. The reason it shifted is
the heart of this post.
Why it matters now
Most of the software you touch as an engineer assumes containers somewhere in the path:
- CI often runs in containers. Many GitHub Actions and GitLab jobs
use a
container:step or container-based executor; even when the outer runner is a VM, build and test steps are routinely wrapped in containers for reproducibility. - Production runs in containers. Kubernetes scheduling, the entire
cloud-native ecosystem, your
Dockerfile— all of it. - Local dev runs in containers. Devcontainers,
docker compose, the Postgres you spin up for a feature branch. - AI/ML serving runs in containers. GPU-enabled containers (via the NVIDIA Container Toolkit) are a common way model servers, training jobs, and inference endpoints get shipped. Model weights are sometimes baked into an image and often mounted or fetched at runtime; either way the runtime is a container with the host’s GPU exposed.
VMs didn’t disappear — they’re still the substrate cloud providers use to isolate tenants from each other on shared hardware, and they show back up in the container world when stronger isolation is needed: Firecracker is a microVM monitor that AWS built for Lambda’s multi-tenant execution environments; Kata launches each container inside a lightweight VM; gVisor takes a different route entirely — no VM at all, just a userspace application kernel that intercepts the container’s syscalls before the host kernel sees them. But for the day-to-day “how do I ship my service” slot, containers have been the dominant answer for years.
The short answer
container = process + namespaces + cgroups + a layered filesystem image
Picture to keep: your Postgres container is not a machine inside your machine — it’s one ordinary process on your host, wearing a blindfold that hides every other process and a filesystem, plus a meter capping how much CPU and RAM it can draw. Where that picture breaks: the blindfold is issued by the host kernel, which is why a kernel bug can lift it, and a VM’s can’t.
A container isn’t a tiny VM. It’s a normal Linux process that the kernel has been told to show a different view of the system to — its own PID 1, its own mount tree, its own network interfaces, its own user IDs — with hard limits on how much CPU and memory it can use. There’s only one kernel: the host’s. That’s why it boots in milliseconds and weighs megabytes.
How it works
The clean way to see it is to build that thirty-second Postgres yourself, starting from the tool the industry already had, and watch each attempt fail into the next.
Attempt 1: give it its own machine
A VM goes deep. The hypervisor emulates a whole computer: virtual CPUs (with help from hardware virtualization extensions like Intel VT-x), virtual RAM, virtual NICs, virtual disks. On top of that emulated hardware, you boot a complete guest operating system — kernel, init, drivers, libc, shell, everything. Your application then runs as a normal process inside that OS.
The isolation is excellent precisely because it’s at the hardware boundary: the guest can’t see the host kernel because it has its own. The cost is also at the hardware boundary: every guest pays for a kernel, memory for that kernel, and the latency of booting it. Your thirty-second Postgres is now a ninety-second Postgres, and ten of them for a test matrix means ten kernels doing nothing but existing.
Fix: drop the second kernel — but then the blindfold has to come from somewhere
What you actually wanted was one Postgres process that can’t see or be seen
by anything else on the box. So run it as a plain process on the host kernel.
Cheap, instant — and broken: it sees every other process, mounts the host’s
/etc and /usr, and binds the host’s port 5432 as itself.
Namespaces are the fix for the seeing half. They scope what a process can
see — and they are necessary rather than sufficient, which is a distinction
worth holding onto until the seams section. namespaces(7)
now lists eight kinds: mount, PID, network, IPC, UTS (hostname), user, cgroup,
and time — the last two arrived well after the set early containers were built
on (cgroup namespaces in 4.6, time namespaces in 5.6). The first process
in a new PID namespace gets PID 1 inside that namespace; from its point of
view, no other processes on the host exist. A process in its own mount
namespace can be given a filesystem rooted somewhere completely different:
the runtime mounts an unpacked image with its own /usr, /lib and so on,
then pivot_roots
into it. The namespace is what keeps that switch private to the container.
Give your
Postgres a network namespace and its port 5432 is its own.
But seeing nothing isn’t the same as taking nothing
A blindfolded process can still eat the machine. Your isolated Postgres runs a runaway query, takes all eight cores and every free page, and every other container on the host starves — you have privacy without fairness.
Cgroups are the fix: they scope what a process can consume. CPU, memory,
block-I/O bandwidth, number of PIDs. Two shapes of knob, and confusing them is
a classic ops mistake: a weight (cgroup v2’s cpu.weight, v1’s “CPU shares”)
only decides who wins when the machine is busy — an idle box lets you use all
of it — while a cap (cpu.max, memory.max) is a ceiling you cannot exceed
no matter how quiet the neighbours are. The kernel enforces both from outside
the container’s view, which is why a container that pushes past its memory cap
and can’t be reclaimed down gets killed rather than politely asked to stop.
That’s the whole isolation story. Run ps -ef on the host while your
Postgres container is running and you’ll see its processes right there in the
host’s process list — regular processes, with extra restrictions on what
they’re allowed to look at and use.
But shipping a root filesystem per service is VM-sized again
The mount namespace needs something to point at: a /usr, a /lib, a libc
of the right version — Postgres’s whole userland. Ship that as a plain
tarball per service and you’re back to gigabyte artifacts and slow pulls,
which was half of what made VMs painful.
The fix is the image format, and it’s the other half of why containers won. A container image is a stack of read-only filesystem layers, each layer a tarball of changed files relative to the layer below, addressed by a content digest of its own bytes (SHA-256, in practice). A typical Python service image might be:
debian:slimbase layer (shared by every Debian-based image)- Python runtime layer (shared by every Python image based on this base)
- Your
pip installlayer (shared by every build with the samerequirements.txt) - Your application code (changes every commit)
When you pull an image, the registry only sends the layers your host doesn’t already have. When you run it, the layers get stacked into a single filesystem by a union filesystem — overlayfs is the usual choice on Linux, though runtimes support other storage backends — with a thin writable layer on top for the running container. Two containers from the same image share the underlying read-only layers in the page cache, so the second one starts even faster than the first.
VM disk formats can do something similar — qcow2 supports backing
files and copy-on-write snapshots, for instance — but the ecosystem
around layered, content-addressed, registry-distributed images
standardized on the container side. In practice, “share a base, only
ship the diff” is what docker pull makes routine, while VM images are
usually shipped as whole filesystems.
And that’s why there’s nothing left to wait for
Put the three fixes together and the “boot” disappears. Starting a container is approximately:
- The higher-level runtime (containerd, which Docker and Kubernetes — via the CRI — typically use) has already prepared the image’s stacked layers as a mountable root filesystem.
- A low-level runtime —
runcis the common one — callsclone()with flags asking for new namespaces. This is a fork-style syscall — see why fork is weird. - It applies the cgroup limits, switches the child’s root to that prepared
filesystem, and
execs your entrypoint binary.
That’s it. No kernel boot, no init system traversal, no driver probing. The first instruction of your application runs almost immediately after the syscall returns. A VM, in contrast, has to POST virtual hardware, run a bootloader, boot a kernel, run an init system, start services, and only then execute your code.
Show the seams
Containers won, but the reasons they didn’t fully replace VMs are worth knowing:
- Same kernel = a shared blast radius. A kernel-level security bug lets a container break out to the host in a way a guest-kernel bug usually can’t, because a VM’s escape still has to get past the hypervisor afterwards. Containers are a boundary, not a weaker VM — the escape targets a different thing. This is why platforms running strangers’ code commonly put a second boundary around it: a lightweight VM (Firecracker, Kata Containers) or, in gVisor’s case, a userspace kernel that answers the container’s syscalls so the host kernel is exposed to far fewer of them.
- Namespaces and cgroups are the skeleton, not the armour. The isolation a
real container runtime gives you also rests on layers this post’s three
fixes don’t include: the kernel capabilities the runtime drops before
exec, a seccomp filter narrowing which syscalls the container may make at all, and an LSM policy (AppArmor or SELinux) on top. Docker, for instance, starts containers with a reduced capability set and a default seccomp profile, and runs user namespaces off by default. “It’s namespaced” is the beginning of a security argument, not the end of one. - No Windows containers on a Linux host (and vice versa). Containers share the host kernel, so the kernel ABI has to match. VMs don’t care.
- GPUs and other devices need explicit pass-through. GPU containers work because the NVIDIA Container Toolkit injects the host’s driver and device files into the container’s mount namespace. It’s not magic; someone wired up the seams.
- “Stateless” is doing real work in this story. Containers are easy to throw away because the design assumes state lives elsewhere (databases, object storage). The moment you put real state inside a container, the layered-image and “cattle, not pets” framing starts fighting you. Kubernetes’ StatefulSets — stable identities, stable storage, ordered rollout — are one answer to that awkwardness, among other things they’re for.
- Cold-start is fast but not free. Pulling a multi-gigabyte AI/ML image (CUDA + PyTorch + model weights) over the network on a cold node is the dominant cost of a “container start” in practice — not the syscalls. Lazy-pull snapshotters (SOCI, stargz) exist to attack this by starting the container before the whole image has landed.
You started with container = process + namespaces + cgroups + a layered filesystem image. What did the chain of fixes add that the line doesn’t say
out loud? — + the host's kernel, shared. That’s the term doing all the work
in both directions: it’s why your Postgres started in under a second (nothing
to boot) and why platforms that run strangers’ code next to yours generally
wrap it in a VM anyway. A VM virtualizes the machine; a
container virtualizes the view from inside one process. Same goal, different
layer — and the container won the deployment slot because its layer was the
right one for “ship my code and its dependencies.”
Check yourself
Before you go — someone benchmarks “container start time” at 4 milliseconds on their laptop, then deploys to a fresh Kubernetes node and sees 90 seconds. Nothing about the container changed. What did?
Answer
The image wasn’t there yet. On the laptop every layer was already in local
storage, so “start” really was just clone() + mount + exec. On a cold node
the runtime has to pull every layer over the network and unpack it before any
of that can happen — and for an AI/ML image with CUDA and PyTorch that’s
multiple gigabytes. (Unless a lazy-pull snapshotter is in play, which is
exactly the point: it starts the container before the whole image has landed.) The millisecond number is real but it measures the last
step only; in production, container start time is usually an image
distribution problem, which is exactly what lazy-pull work like SOCI and
stargz attacks.
And one more — if containers are so much lighter, why did AWS build Firecracker and wrap Lambda’s execution environments in microVMs instead?
Answer
Because the isolation boundary, not the weight, is what Lambda needs. Lambda runs arbitrary code from strangers, side by side, on shared hardware. Container isolation is enforced by the host kernel, so one kernel bug is one escape away from every tenant on that box. A VM moves the boundary: compromising the guest kernel gets the attacker root in a machine that is itself a sandbox, and reaching other tenants means getting through the hypervisor as well — a much smaller, much more heavily scrutinised surface. Firecracker’s bet is that you can shrink a VM (minimal device model, no BIOS, no legacy hardware) until it boots fast enough to compete with containers on startup while keeping the stronger boundary. It’s the “same kernel = shared blast radius” seam above, priced.
Famous related terms
- Namespace —
namespace = a per-process kernel-level scope for some kind of resource (filesystem, PID, network…)— the “what can I see” half of containerization. - cgroup —
cgroup = kernel-enforced quota and accounting for a group of processes— the “what can I use” half. - Hypervisor —
hypervisor = thin layer that virtualizes hardware so multiple guest OSes can run on one machine— what a VM runs on. - Image layer —
layer = tarball of file changes + content-addressed hash— whydocker pullis incremental and image storage deduplicates. - OCI —
OCI = Open Container Initiative ≈ vendor-neutral specs for container images and runtimes— whatdocker buildproduces is an OCI-compatible image; OCI runtimes likeruncthen run an unpacked OCI bundle derived from it. - Firecracker —
Firecracker = minimal VMM + microVMs that boot in milliseconds— what AWS Lambda and Fargate use to get VM-grade isolation at container-grade speed. - Kata Containers —
Kata = OCI runtime + each container in its own lightweight VM— container ergonomics, VM isolation boundary. - gVisor —
gVisor = userspace application kernel that intercepts syscalls + OCI runtime— a different bet: keep one host kernel, but have container syscalls hit a sandbox kernel first.
Going deeper
- Linux man pages:
namespaces(7),cgroups(7),clone(2)— the primary source for “what exactly does the kernel give you,” i.e. the real list of namespace types and cgroup controllers behind every runtime. - Jérôme Petazzoni’s talk “Cgroups, namespaces and beyond: what are containers made from?” — the explainer to watch if you want someone to build a container in front of you from raw syscalls, with no Docker layer in the way.
- The OCI image-spec and runtime-spec — the rabbit hole: what an image and a “running container” are contractually, once you stop assuming Docker is the implementation.