Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why containers won over VMs

Both promise isolated, reproducible environments. One boots in milliseconds and ships in megabytes; the other boots in seconds and ships in gigabytes. The reason isn't 'containers are lighter VMs' — they're a different kind of thing entirely.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has typed docker run but never asked what it actually did. The prose below fills in the seams.

1 · The problem

The same request, sixteen years apart.

2010 download a GB image boot a kernel start an init system finally, Postgres today docker run postgres — and it is accepting connections Same isolation goal. No kernel to boot.
What changed is not that virtual machines got slow. It is that somebody found a different layer to put the boundary at.

2 · The old way

A VM duplicates the whole machine, kernel and all.

a virtual machine your Postgres guest userland guest kernel hypervisor host kernel hardware two whole layers your workload never uses — and ten copies for a test matrix means ten kernels doing nothing what you actually wanted my code, my dependencies, and nobody else’s stuff in sight
Nobody wanted a second kernel. They wanted isolation of the things above the kernel — and a VM is the wrong instrument for that, not a badly built one.

3 · The trick

A blindfold and a meter, issued to one ordinary process.

postgres one process running on the host kernel blindfold meter NAMESPACES — what it can see its own PID 1, its own mount tree, its own network interfaces. Every other process is invisible. CGROUPS — what it can use a weight decides who wins when the box is busy. a cap is a ceiling, however quiet the neighbours are.
Run ps on the host and the container’s processes are right there in the list — ordinary processes, with extra restrictions on what they may look at and use.

4 · The other half of the win

The image is a stack, so you only ship the top of it.

one image = a stack of layers what docker pull sends your application code the only thing shipped your pip install already on the host the Python runtime already on the host debian:slim base already on the host The registry sends only what you don’t have. and two containers from one image share those layers in the page cache
Each layer is named by a hash of its own bytes, so identical content is identical everywhere. “Share a base, ship the diff” is what turns a gigabyte artifact into a fast pull.

5 · The seam

The blindfold is issued by the host’s own kernel.

container container container ONE HOST KERNEL a bug here reaches all of them container guest kernel hypervisor an escape must get past this too A different boundary, not a weaker VM. which is why capabilities, seccomp and an LSM policy sit on top of every real runtime
Sharing the kernel is what makes a container instant and what makes the blast radius shared. Platforms that run strangers’ code next to yours generally wrap it in a VM anyway.

6 · Keep this card

The whole thing on one index card.

CONTAINER = one ordinary process + a blindfold — namespaces + a meter — cgroups + a stack of shared layers …all of it issued by the host’s own kernel
Picture to keep: not a machine inside your machine — one ordinary process wearing a blindfold that hides everyone else, with a meter capping what it can draw. A VM virtualizes the machine; a container virtualizes the view from inside one process.

Why it exists

You need a Postgres for a feature branch. You type docker run postgres, and before you’ve finished alt-tabbing back to your editor it’s accepting connections. When you’re done you delete it and the machine looks like it never happened. That thirty-second Postgres is the running example for this post — keep it in mind, because the same request in 2010 meant downloading a multi-gigabyte disk image, booting a whole operating system inside your computer, and waiting.

That’s what deploying a service used to mean: a virtual machine — a full guest operating system, kernel and all, running on top of a hypervisor on top of the host kernel. It worked, but it was heavy. A “small” service shipped as a multi-gigabyte disk image, took tens of seconds to boot, and spent most of its RAM on a kernel and userland its actual workload would never use. Running ten copies of your service for testing meant ten kernels.

The pain point was a mismatch. What developers actually wanted was: “give me my code, my dependencies, and a filesystem that looks the way I expect — and please don’t let me see anyone else’s stuff.” They didn’t want a second kernel. They wanted isolation of the things above the kernel, not duplication of the kernel.

Linux had been quietly accumulating the pieces to do exactly that: namespaces (starting with mount namespaces in 2002; PID, network and the rest arrived over the following decade, and user namespaces landed in 3.8 / 2013 — later hardening rather than something the first containers were built on) and cgroups (started at Google and shipped in Linux 2.6.24, January 2008). In 2013, Docker packaged those primitives behind a friendly CLI and an image format you could push and pull, and over the next several years much of the industry’s deployment story shifted from VMs to containers. The reason it shifted is the heart of this post.

Why it matters now

Most of the software you touch as an engineer assumes containers somewhere in the path:

VMs didn’t disappear — they’re still the substrate cloud providers use to isolate tenants from each other on shared hardware, and they show back up in the container world when stronger isolation is needed: Firecracker is a microVM monitor that AWS built for Lambda’s multi-tenant execution environments; Kata launches each container inside a lightweight VM; gVisor takes a different route entirely — no VM at all, just a userspace application kernel that intercepts the container’s syscalls before the host kernel sees them. But for the day-to-day “how do I ship my service” slot, containers have been the dominant answer for years.

The short answer

container = process + namespaces + cgroups + a layered filesystem image

Picture to keep: your Postgres container is not a machine inside your machine — it’s one ordinary process on your host, wearing a blindfold that hides every other process and a filesystem, plus a meter capping how much CPU and RAM it can draw. Where that picture breaks: the blindfold is issued by the host kernel, which is why a kernel bug can lift it, and a VM’s can’t.

A container isn’t a tiny VM. It’s a normal Linux process that the kernel has been told to show a different view of the system to — its own PID 1, its own mount tree, its own network interfaces, its own user IDs — with hard limits on how much CPU and memory it can use. There’s only one kernel: the host’s. That’s why it boots in milliseconds and weighs megabytes.

How it works

The clean way to see it is to build that thirty-second Postgres yourself, starting from the tool the industry already had, and watch each attempt fail into the next.

Attempt 1: give it its own machine

A VM goes deep. The hypervisor emulates a whole computer: virtual CPUs (with help from hardware virtualization extensions like Intel VT-x), virtual RAM, virtual NICs, virtual disks. On top of that emulated hardware, you boot a complete guest operating system — kernel, init, drivers, libc, shell, everything. Your application then runs as a normal process inside that OS.

The isolation is excellent precisely because it’s at the hardware boundary: the guest can’t see the host kernel because it has its own. The cost is also at the hardware boundary: every guest pays for a kernel, memory for that kernel, and the latency of booting it. Your thirty-second Postgres is now a ninety-second Postgres, and ten of them for a test matrix means ten kernels doing nothing but existing.

Fix: drop the second kernel — but then the blindfold has to come from somewhere

What you actually wanted was one Postgres process that can’t see or be seen by anything else on the box. So run it as a plain process on the host kernel. Cheap, instant — and broken: it sees every other process, mounts the host’s /etc and /usr, and binds the host’s port 5432 as itself.

Namespaces are the fix for the seeing half. They scope what a process can see — and they are necessary rather than sufficient, which is a distinction worth holding onto until the seams section. namespaces(7) now lists eight kinds: mount, PID, network, IPC, UTS (hostname), user, cgroup, and time — the last two arrived well after the set early containers were built on (cgroup namespaces in 4.6, time namespaces in 5.6). The first process in a new PID namespace gets PID 1 inside that namespace; from its point of view, no other processes on the host exist. A process in its own mount namespace can be given a filesystem rooted somewhere completely different: the runtime mounts an unpacked image with its own /usr, /lib and so on, then pivot_roots into it. The namespace is what keeps that switch private to the container. Give your Postgres a network namespace and its port 5432 is its own.

But seeing nothing isn’t the same as taking nothing

A blindfolded process can still eat the machine. Your isolated Postgres runs a runaway query, takes all eight cores and every free page, and every other container on the host starves — you have privacy without fairness.

Cgroups are the fix: they scope what a process can consume. CPU, memory, block-I/O bandwidth, number of PIDs. Two shapes of knob, and confusing them is a classic ops mistake: a weight (cgroup v2’s cpu.weight, v1’s “CPU shares”) only decides who wins when the machine is busy — an idle box lets you use all of it — while a cap (cpu.max, memory.max) is a ceiling you cannot exceed no matter how quiet the neighbours are. The kernel enforces both from outside the container’s view, which is why a container that pushes past its memory cap and can’t be reclaimed down gets killed rather than politely asked to stop.

That’s the whole isolation story. Run ps -ef on the host while your Postgres container is running and you’ll see its processes right there in the host’s process list — regular processes, with extra restrictions on what they’re allowed to look at and use.

But shipping a root filesystem per service is VM-sized again

The mount namespace needs something to point at: a /usr, a /lib, a libc of the right version — Postgres’s whole userland. Ship that as a plain tarball per service and you’re back to gigabyte artifacts and slow pulls, which was half of what made VMs painful.

The fix is the image format, and it’s the other half of why containers won. A container image is a stack of read-only filesystem layers, each layer a tarball of changed files relative to the layer below, addressed by a content digest of its own bytes (SHA-256, in practice). A typical Python service image might be:

When you pull an image, the registry only sends the layers your host doesn’t already have. When you run it, the layers get stacked into a single filesystem by a union filesystem — overlayfs is the usual choice on Linux, though runtimes support other storage backends — with a thin writable layer on top for the running container. Two containers from the same image share the underlying read-only layers in the page cache, so the second one starts even faster than the first.

VM disk formats can do something similar — qcow2 supports backing files and copy-on-write snapshots, for instance — but the ecosystem around layered, content-addressed, registry-distributed images standardized on the container side. In practice, “share a base, only ship the diff” is what docker pull makes routine, while VM images are usually shipped as whole filesystems.

And that’s why there’s nothing left to wait for

Put the three fixes together and the “boot” disappears. Starting a container is approximately:

  1. The higher-level runtime (containerd, which Docker and Kubernetes — via the CRI — typically use) has already prepared the image’s stacked layers as a mountable root filesystem.
  2. A low-level runtime — runc is the common one — calls clone() with flags asking for new namespaces. This is a fork-style syscall — see why fork is weird.
  3. It applies the cgroup limits, switches the child’s root to that prepared filesystem, and execs your entrypoint binary.

That’s it. No kernel boot, no init system traversal, no driver probing. The first instruction of your application runs almost immediately after the syscall returns. A VM, in contrast, has to POST virtual hardware, run a bootloader, boot a kernel, run an init system, start services, and only then execute your code.

Show the seams

Containers won, but the reasons they didn’t fully replace VMs are worth knowing:

You started with container = process + namespaces + cgroups + a layered filesystem image. What did the chain of fixes add that the line doesn’t say out loud? — + the host's kernel, shared. That’s the term doing all the work in both directions: it’s why your Postgres started in under a second (nothing to boot) and why platforms that run strangers’ code next to yours generally wrap it in a VM anyway. A VM virtualizes the machine; a container virtualizes the view from inside one process. Same goal, different layer — and the container won the deployment slot because its layer was the right one for “ship my code and its dependencies.”

Check yourself

Before you go — someone benchmarks “container start time” at 4 milliseconds on their laptop, then deploys to a fresh Kubernetes node and sees 90 seconds. Nothing about the container changed. What did?

Answer

The image wasn’t there yet. On the laptop every layer was already in local storage, so “start” really was just clone() + mount + exec. On a cold node the runtime has to pull every layer over the network and unpack it before any of that can happen — and for an AI/ML image with CUDA and PyTorch that’s multiple gigabytes. (Unless a lazy-pull snapshotter is in play, which is exactly the point: it starts the container before the whole image has landed.) The millisecond number is real but it measures the last step only; in production, container start time is usually an image distribution problem, which is exactly what lazy-pull work like SOCI and stargz attacks.

And one more — if containers are so much lighter, why did AWS build Firecracker and wrap Lambda’s execution environments in microVMs instead?

Answer

Because the isolation boundary, not the weight, is what Lambda needs. Lambda runs arbitrary code from strangers, side by side, on shared hardware. Container isolation is enforced by the host kernel, so one kernel bug is one escape away from every tenant on that box. A VM moves the boundary: compromising the guest kernel gets the attacker root in a machine that is itself a sandbox, and reaching other tenants means getting through the hypervisor as well — a much smaller, much more heavily scrutinised surface. Firecracker’s bet is that you can shrink a VM (minimal device model, no BIOS, no legacy hardware) until it boots fast enough to compete with containers on startup while keeping the stronger boundary. It’s the “same kernel = shared blast radius” seam above, priced.

Going deeper