Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why virtual memory exists

Every process thinks it owns the whole machine. That lie is the foundation almost every modern OS feature is built on.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never wondered why the numbers in Activity Monitor don’t add up. The prose below fills in the seams.

1 · The problem

The memory column adds up to more RAM than you own.

Memory browser, 40 tabs 9.4 GB dev server 3.6 GB editor 2.1 GB Slack 1.8 GB Spotify 0.9 GB total 17.8 GB > 16 GB of actual RAM and nothing crashed the machine feels fine
Ask each program how much it has allocated and the total is wilder still. The reason the arithmetic works anyway is virtual memory.

2 · Take it away

Without the trick, two programs want the same byte.

program A wants 0x1000 program B wants 0x1000 burned into the binary PHYSICAL RAM 0x1000 there is exactly one of it They collide. One scribbles over the other. and running your forty-tab browser is not merely slow — it is unwritable
No isolation, no swap, no program bigger than RAM, and no two programs at once without the authors agreeing on addresses in advance. All four problems are the same problem.

3 · The lie

Hand every process the same map. Translate at the door.

process A 0 … the whole map process B 0 … the whole map process C 0 … the whole map identical fake addresses TRANSLATOR page table + MMU hardware, at the door or refuses REAL RAM a different page for each of them Every process owns the whole machine. the hardware only says yes or refuses; every real decision is the kernel’s
The addresses are fake, so they can be handed out freely: isolation, shared libraries mapped once into many processes, and ASLR all fall straight out of the renaming.

4 · Making the map affordable

Two fixes, because the obvious table is bigger than the memory.

one entry per byte bigger than RAM the table describing memory exceeds it absurd fix one entry per 4 KB page 512 GB per process a flat array covering a 48-bit address space still absurd fix a tree, full of holes a few hundred KB branches for untouched ranges never exist affordable but walking a tree costs four extra loads per access so the CPU caches recent translations — the TLB — and most accesses hit it
Each fix is forced by the previous one’s failure. That chain is the whole design: pages, then a sparse tree, then a cache in front of the tree.

5 · Where the power actually is

The translation is allowed to fail on purpose.

the translator refuses → the kernel gets a turn PAGE FAULT five different meanings lazy allocation first touch of new memory → make a page, restart file page-in the page belongs to a mapped file → read it, restart swap-in it was evicted to disk → read it back, restart copy-on-write shared read-only, and someone wrote → copy, restart segfault that address was never yours → SIGSEGV the first four are invisible to your program — only the last one makes the news
Swap, mmap, lazy allocation, copy-on-write and the OOM killer are all the same move: arrange for the hardware to refuse at a useful moment, then do the real work in the handler.

6 · Keep this card

The whole thing on one index card.

VIRTUAL MEMORY = a fake address space per process + a table mapping fake → real + permission to refuse …the renaming buys isolation; the refusal buys the rest
Picture to keep: every process is handed its own copy of the same street map with the same house numbers, and a translator at the door quietly rewrites each address into a real location — or says “that house isn’t built yet, hold on” and builds it while you wait.

Why it exists

Open Activity Monitor (or Task Manager) on a 16 GB laptop while you’re doing normal work: a browser with forty tabs, Slack, an editor, a dev server, Spotify. Add up the memory column. It can easily come to more than 16 GB — and yet nothing has crashed, nothing is complaining, and the machine feels fine. Ask each process how much memory it has allocated and the total is wilder still. That laptop is the running example for this post, and the reason the numbers don’t add up is virtual memory.

To see why anyone built it, take it away. Picture a machine where the linker burns the absolute physical addresses of every variable and function straight into the binary. The program only runs if those exact bytes of RAM are free. Run two programs at once and they fight over the same addresses. Crash one and it can scribble all over the other. Want to use more memory than the box has? You do it by hand — machines of that era had overlays, base/limit relocation registers, and programmer-managed swapping, so it wasn’t literally impossible, just your job. Your forty-tab browser is not merely slow in that world; nobody could write it.

The pain points stack up fast:

Virtual memory solves all four with one move: lie to every process. Each process gets its own private, contiguous-looking address space, starting from zero, as if it owned the entire machine. Underneath, the MMU translates those fake addresses to real physical RAM — or refuses, which hands control to the kernel, which then decides what the refusal meant: “it’s out on disk,” “it doesn’t exist yet, allocate it on first touch,” “it’s shared with another process and you just tried to write it.” The hardware only ever says yes or no; every interesting answer is the kernel’s.

Once you have the indirection, the features stack up on top of it: process isolation, shared libraries mapped once into many processes, ASLR. Once the translation is also allowed to refuse, you get the rest: overcommit, copy-on-write fork, memory-mapped files, swap, lazy allocation. None of it is possible without the lie.

Why it matters now

It’s invisible until it isn’t. Every time you:

…you are leaning on virtual memory — and so is the laptop above, every time you click a tab you haven’t looked at since this morning. Container memory limits, cgroup accounting, Postgres MVCC sharing buffer pages between processes, GPU drivers mapping VRAM into your address space — all of it sits on top of the same trick.

The AI-era version: model weights and KV caches are huge, and a lot of the “how do I fit this on this box” art is really how do I cooperate with virtual memory — using mmap so the OS pages weights in lazily, sharing read-only pages between worker processes, or going the other way and pinning pages so the GPU can DMA them.

The short answer

virtual memory = per-process fake address space + page table that maps fake → real, or refuses so the kernel can decide

Picture to keep: every process is handed its own copy of the same street map with the same house numbers, and a translator at the door quietly rewrites each address into a real location — or says “that house isn’t built yet, hold on” and builds it while you wait. Where the analogy breaks: the translator is hardware and can only say yes or refuse. Every interesting decision — build it, fetch it back from disk, or crash you — happens after the refusal, in the kernel.

Every process sees its own address space. Every memory access is translated through a per-process page table — usually straight from a cached translation, and only by actually walking the table when the cache misses. If the entry can’t satisfy the access, the kernel gets a chance to fix it up: load from disk, allocate a new page, copy a shared one, or kill the process for touching memory it doesn’t own. Note the lossy part of the compression line: the hardware entry points at physical memory or is unusable — “this page is out on swap” is the kernel’s bookkeeping, not a field the MMU reads.

How it works

The best way to see the design is to try to build it yourself for that laptop, and watch each attempt break.

Attempt 1: give every process fake addresses and keep a translation table. The idea is right — that’s the whole trick — but be naive about the table and it collapses immediately. One entry per byte means the translation table is larger than the memory it describes.

Fix: translate in chunks. Memory is divided into fixed-size pages — typically 4 KB on x86-64. The unit of mapping is the page, not the byte, so one table entry now covers 4,096 addresses. A browser tab’s 200 MB of heap needs ~50,000 entries instead of 200 million.

But the table is still absurd. x86-64’s classic 4-level paging covers a 48-bit virtual address space. At 4 KB per page that’s about 68 billion pages, and a flat array of 8-byte entries to describe them would cost 512 GB — per process. Your forty tabs are each using a sliver of that space and leaving vast holes untouched.

Fix: make the table a tree. Each process’s page table is a tree whose leaves are PTEs, each saying “virtual page N maps to physical page M, with these permissions (read/write/execute), and here’s a ‘present’ bit.” Branches for address ranges nobody touched simply don’t exist. On x86-64 it’s a 4-level tree (5 levels on newer chips supporting a 57-bit address space). A process using a few hundred MB pays for a few hundred KB of tables, not 512 GB.

But now every memory access costs four extra memory accesses. To translate one address, the MMU must walk the tree from the root — that’s four dependent loads before your actual load happens. Doing that on every instruction would make the machine several times slower.

Fix: cache the translations. The TLB is a small hardware cache of recent virtual→physical translations. Hits are essentially free; only misses pay for the page walk. Because programs touch the same pages over and over, most accesses hit. This is also why “huge pages” of 2 MB or 1 GB exist — one entry covers far more memory, so a big-heap workload misses less.

And the original problem — more memory than the box has — is still unsolved. Everything so far just renames addresses. The fix is what makes virtual memory powerful rather than merely tidy: a translation is allowed to fail on purpose. When the MMU hits an entry it can’t use for this access — the present bit is clear, or the page is there but the permissions say no — it doesn’t return garbage and it doesn’t halt. It raises a page fault, which hands control to the kernel, and the kernel gets to decide what that fault means:

  1. Lazy allocation. First write to a freshly malloc’d page. Allocate a real page, zero it, fix up the PTE, restart the instruction. The program never knows. (Linux is stingier still: a first read of untouched anonymous memory is often satisfied by pointing the entry at one shared zero page, and the private allocation waits for a write.)
  2. File-backed page-in. The page belongs to a memory-mapped file. Read it from disk, install it, restart.
  3. Swap-in. The page was evicted to swap. Read it back, install it, restart.
  4. Copy-on-write. Two processes share a page read-only after fork. One tries to write. Note that the page is present here — the fault comes from the permission bits, not the present bit. The kernel allocates a fresh copy, points the writer’s PTE at it, restarts. Until that moment, fork was nearly free.
  5. Segfault. The address isn’t mapped at all, or the access violates permissions (write to read-only, execute on no-execute). Kernel sends SIGSEGV. The famous segfault is just “the page table said no.”

Cases 1–4 are invisible to the program. Case 5 is the only one users see, and even then only as “the program crashed.”

This is what makes the laptop’s arithmetic work. A background tab untouched for an hour can have its pages compressed or evicted to disk and faulted back when you click it. A freshly malloc’d buffer nobody has written to yet has no physical page behind it at all — case 1 will conjure one on first touch. The libc pages every process maps exist once in RAM but get counted against each process that maps them. So the column you added up is double-counting shared pages and describing memory that partly isn’t resident; the page fault is the machinery that keeps the difference invisible.

What you actually get from the indirection

Once translation exists, layering features onto it is almost free:

Show the seams

The deepest seam: virtual memory is the OS pretending it has more control than the hardware really gives it. The MMU is hardware; the kernel only gets a turn when the MMU faults. So a surprising number of “OS features” in this space are really “arrange for the MMU to fault at a useful moment, then do the work in the handler.” Page-on-demand, copy-on-write, swap, mmap — all the same shape, and some GC and JIT tricks borrow it too (a collector can protect a region so writes to it trap).

So: the laptop’s memory column doesn’t add up because most of what it lists isn’t a claim on physical RAM — it’s an address, and an address is free. You started with virtual memory = per-process fake address space + page table. What did the chain of fixes add? — + a translation that's allowed to fail. Sort the features by which half of the line pays for them and the design gets sharp. The renaming alone buys the static wins — isolation, shared libraries, ASLR — because all three are just “put this process’s pages at addresses of my choosing.” The refusal buys everything dynamic: swap, mmap, lazy allocation, copy-on-write, and the OOM killer, because every one of those is the kernel doing work inside a fault it arranged to happen.

Check yourself

Before you go — a colleague says their service “leaks memory,” pointing at a virtual-size number of 40 GB on a 32 GB box. What would you ask to see before agreeing?

Answer

Resident size, not virtual size. Virtual size counts address space the process has reserved — including mmaped files it barely reads, guard regions, and allocations never touched. None of that consumes physical RAM until a page fault forces the kernel to back it. A process can have a 40 GB virtual size and a 300 MB resident set and be perfectly healthy. The number that tells you about a leak is resident memory climbing over time (plus swap-in/swap-out rates and whether the box is thrashing).

And one more — you fork() a 10 GB process and the child immediately execs a tiny binary. Almost nothing was copied. Now instead the child stays and starts writing all over its inherited heap. What happens to your memory usage, and why?

Answer

It climbs toward a second 10 GB. fork was cheap because parent and child shared every page read-only; the copy only happens per-page, on the first write. An exec throws the whole inherited address space away before those writes ever happen, so the copies never occur. A child that stays and writes is walking through case 4 (copy-on-write) page by page — and each one costs a fault plus a real 4 KB allocation. “Fork is cheap” is a statement about deferred work, not about work that never happens.

Going deeper