Why virtual memory exists
Every process thinks it owns the whole machine. That lie is the foundation almost every modern OS feature is built on.
On this page
The picture version
The whole idea in six pictures, for a reader who has never wondered why the numbers in Activity Monitor don’t add up. The prose below fills in the seams.
1 · The problem
The memory column adds up to more RAM than you own.
2 · Take it away
Without the trick, two programs want the same byte.
3 · The lie
Hand every process the same map. Translate at the door.
4 · Making the map affordable
Two fixes, because the obvious table is bigger than the memory.
5 · Where the power actually is
The translation is allowed to fail on purpose.
6 · Keep this card
The whole thing on one index card.
Why it exists
Open Activity Monitor (or Task Manager) on a 16 GB laptop while you’re doing normal work: a browser with forty tabs, Slack, an editor, a dev server, Spotify. Add up the memory column. It can easily come to more than 16 GB — and yet nothing has crashed, nothing is complaining, and the machine feels fine. Ask each process how much memory it has allocated and the total is wilder still. That laptop is the running example for this post, and the reason the numbers don’t add up is virtual memory.
To see why anyone built it, take it away. Picture a machine where the linker burns the absolute physical addresses of every variable and function straight into the binary. The program only runs if those exact bytes of RAM are free. Run two programs at once and they fight over the same addresses. Crash one and it can scribble all over the other. Want to use more memory than the box has? You do it by hand — machines of that era had overlays, base/limit relocation registers, and programmer-managed swapping, so it wasn’t literally impossible, just your job. Your forty-tab browser is not merely slow in that world; nobody could write it.
The pain points stack up fast:
- Programs can’t coexist without coordinating addresses, which they can’t.
- No isolation — any pointer bug in any program can corrupt any other.
- No way to use a disk as overflow — addresses are physical, and disks aren’t memory.
- No way to run a program bigger than RAM without the programmer hand-rolling it, one overlay at a time.
Virtual memory solves all four with one move: lie to every process. Each process gets its own private, contiguous-looking address space, starting from zero, as if it owned the entire machine. Underneath, the MMU translates those fake addresses to real physical RAM — or refuses, which hands control to the kernel, which then decides what the refusal meant: “it’s out on disk,” “it doesn’t exist yet, allocate it on first touch,” “it’s shared with another process and you just tried to write it.” The hardware only ever says yes or no; every interesting answer is the kernel’s.
Once you have the indirection, the features stack up on top of it: process
isolation, shared libraries mapped once into many processes,
ASLR.
Once the translation is also allowed to refuse, you get the rest:
overcommit,
copy-on-write fork, memory-mapped files, swap, lazy allocation. None of it
is possible without the lie.
Why it matters now
It’s invisible until it isn’t. Every time you:
- run more processes than fit in RAM,
mmapa 200 GB file and read it like an array,- load a 70 GB LLM’s weights from disk and let the OS page them in,
- watch your dev server’s memory go up but your machine stay snappy,
- get killed by the OOM killer for touching pages the kernel had only promised you,
…you are leaning on virtual memory — and so is the laptop above, every time you
click a tab you haven’t looked at since this morning. Container memory limits, cgroup
accounting, Postgres MVCC sharing buffer pages
between processes, GPU drivers mapping VRAM into your address space — all of
it sits on top of the same trick.
The AI-era version: model weights and KV caches are huge, and a lot of the
“how do I fit this on this box” art is really how do I cooperate with virtual
memory — using mmap so the OS pages weights in lazily, sharing read-only
pages between worker processes, or going the other way and pinning pages so
the GPU can DMA them.
The short answer
virtual memory = per-process fake address space + page table that maps fake → real, or refuses so the kernel can decide
Picture to keep: every process is handed its own copy of the same street map with the same house numbers, and a translator at the door quietly rewrites each address into a real location — or says “that house isn’t built yet, hold on” and builds it while you wait. Where the analogy breaks: the translator is hardware and can only say yes or refuse. Every interesting decision — build it, fetch it back from disk, or crash you — happens after the refusal, in the kernel.
Every process sees its own address space. Every memory access is translated through a per-process page table — usually straight from a cached translation, and only by actually walking the table when the cache misses. If the entry can’t satisfy the access, the kernel gets a chance to fix it up: load from disk, allocate a new page, copy a shared one, or kill the process for touching memory it doesn’t own. Note the lossy part of the compression line: the hardware entry points at physical memory or is unusable — “this page is out on swap” is the kernel’s bookkeeping, not a field the MMU reads.
How it works
The best way to see the design is to try to build it yourself for that laptop, and watch each attempt break.
Attempt 1: give every process fake addresses and keep a translation table. The idea is right — that’s the whole trick — but be naive about the table and it collapses immediately. One entry per byte means the translation table is larger than the memory it describes.
Fix: translate in chunks. Memory is divided into fixed-size pages — typically 4 KB on x86-64. The unit of mapping is the page, not the byte, so one table entry now covers 4,096 addresses. A browser tab’s 200 MB of heap needs ~50,000 entries instead of 200 million.
But the table is still absurd. x86-64’s classic 4-level paging covers a 48-bit virtual address space. At 4 KB per page that’s about 68 billion pages, and a flat array of 8-byte entries to describe them would cost 512 GB — per process. Your forty tabs are each using a sliver of that space and leaving vast holes untouched.
Fix: make the table a tree. Each process’s page table is a tree whose leaves are PTEs, each saying “virtual page N maps to physical page M, with these permissions (read/write/execute), and here’s a ‘present’ bit.” Branches for address ranges nobody touched simply don’t exist. On x86-64 it’s a 4-level tree (5 levels on newer chips supporting a 57-bit address space). A process using a few hundred MB pays for a few hundred KB of tables, not 512 GB.
But now every memory access costs four extra memory accesses. To translate one address, the MMU must walk the tree from the root — that’s four dependent loads before your actual load happens. Doing that on every instruction would make the machine several times slower.
Fix: cache the translations. The TLB is a small hardware cache of recent virtual→physical translations. Hits are essentially free; only misses pay for the page walk. Because programs touch the same pages over and over, most accesses hit. This is also why “huge pages” of 2 MB or 1 GB exist — one entry covers far more memory, so a big-heap workload misses less.
And the original problem — more memory than the box has — is still unsolved. Everything so far just renames addresses. The fix is what makes virtual memory powerful rather than merely tidy: a translation is allowed to fail on purpose. When the MMU hits an entry it can’t use for this access — the present bit is clear, or the page is there but the permissions say no — it doesn’t return garbage and it doesn’t halt. It raises a page fault, which hands control to the kernel, and the kernel gets to decide what that fault means:
- Lazy allocation. First write to a freshly
malloc’d page. Allocate a real page, zero it, fix up the PTE, restart the instruction. The program never knows. (Linux is stingier still: a first read of untouched anonymous memory is often satisfied by pointing the entry at one shared zero page, and the private allocation waits for a write.) - File-backed page-in. The page belongs to a memory-mapped file. Read it from disk, install it, restart.
- Swap-in. The page was evicted to swap. Read it back, install it, restart.
- Copy-on-write. Two processes share a page read-only after
fork. One tries to write. Note that the page is present here — the fault comes from the permission bits, not the present bit. The kernel allocates a fresh copy, points the writer’s PTE at it, restarts. Until that moment,forkwas nearly free. - Segfault. The address isn’t mapped at all, or the access violates
permissions (write to read-only, execute on no-execute). Kernel sends
SIGSEGV. The famous segfault is just “the page table said no.”
Cases 1–4 are invisible to the program. Case 5 is the only one users see, and even then only as “the program crashed.”
This is what makes the laptop’s arithmetic work. A background tab untouched for
an hour can have its pages compressed or evicted to disk and faulted back when
you click it. A freshly malloc’d buffer nobody has written to yet has no
physical page behind it at all — case 1 will conjure one on first touch. The
libc pages every process maps exist once in RAM but get counted against each
process that maps them. So the column you added up is double-counting shared
pages and describing memory that partly isn’t resident; the page fault is the
machinery that keeps the difference invisible.
What you actually get from the indirection
Once translation exists, layering features onto it is almost free:
- Isolation — process A simply has no PTE pointing at process B’s pages.
- Shared libraries —
libc’s code pages exist once in RAM, mapped read-only into every process that uses it. mmap— treat a file as an array; the kernel pages it in on demand.- Overcommit — hand out virtual pages without backing them with physical
ones until they’re touched. On Linux, with its default policy, this is where
the OOM killer comes from; systems that refuse to overcommit fail your
mallocup front instead. - ASLR — pick different virtual addresses each run so attackers can’t predict where things land. Cheap because the addresses are fake anyway.
Show the seams
- It is not free. Every memory access pays for the page walk if the TLB misses. Workloads with terrible locality (some hash tables, some graph traversals) can spend a real fraction of their time in TLB misses. Huge pages exist mostly to amortize this.
- Page tables themselves consume memory. A 4 KB page on x86-64 needs a PTE somewhere; covering a terabyte of address space costs measurable RAM for the tables alone. This is part of why huge pages matter for big-memory workloads.
- Fork + exec is doing more than it looks. A
fork()of a 10 GB process doesn’t copy 10 GB; it duplicates page tables and marks pages copy-on-write. A typicalfork-then-execpattern barely touches anything before tossing the child’s address space away. This is why UNIX’s “fork is cheap” claim survives. - Swap is not “more RAM.” It’s “RAM with disk latency on miss.” Once the working set exceeds RAM, the system thrashes — page-fault, page-in, evict, page-fault again — and effective throughput collapses. The invisibility of paging is also its trap.
- The exact PTE format is architecture-specific. I’ve described x86-64; ARM64 differs in details (and uses two separate page-table bases for user/kernel). The shape — multi-level tree, TLB, faults — is universal, but the bits aren’t.
The deepest seam: virtual memory is the OS pretending it has more control than
the hardware really gives it. The MMU is hardware; the kernel only gets a
turn when the MMU faults. So a surprising number of “OS features” in this space
are really “arrange for the MMU to fault at a useful moment, then do the work
in the handler.” Page-on-demand, copy-on-write, swap, mmap — all the same
shape, and some GC and JIT tricks borrow it too (a collector can protect a
region so writes to it trap).
So: the laptop’s memory column doesn’t add up because most of what it lists
isn’t a claim on physical RAM — it’s an address, and an address is free. You
started with virtual memory = per-process fake address space + page table.
What did the chain of fixes add? — + a translation that's allowed to fail.
Sort the features by which half of the line pays for them and the design gets
sharp. The renaming alone buys the static wins — isolation, shared libraries,
ASLR — because all three are just “put this process’s pages at addresses of my
choosing.” The refusal buys everything dynamic: swap, mmap, lazy
allocation, copy-on-write, and the OOM killer, because every one of those is
the kernel doing work inside a fault it arranged to happen.
Check yourself
Before you go — a colleague says their service “leaks memory,” pointing at a virtual-size number of 40 GB on a 32 GB box. What would you ask to see before agreeing?
Answer
Resident size, not virtual size. Virtual size counts address space the process
has reserved — including mmaped files it barely reads, guard regions, and
allocations never touched. None of that consumes physical RAM until a page fault
forces the kernel to back it. A process can have a 40 GB virtual size and a
300 MB resident set and be perfectly healthy. The number that tells you about a
leak is resident memory climbing over time (plus swap-in/swap-out rates and
whether the box is thrashing).
And one more — you fork() a 10 GB process and the child immediately execs a
tiny binary. Almost nothing was copied. Now instead the child stays and starts
writing all over its inherited heap. What happens to your memory usage, and why?
Answer
It climbs toward a second 10 GB. fork was cheap because parent and child
shared every page read-only; the copy only happens per-page, on the first write.
An exec throws the whole inherited address space away before those writes ever
happen, so the copies never occur. A child that stays and writes is walking
through case 4 (copy-on-write) page by page — and each one costs a fault plus a
real 4 KB allocation. “Fork is cheap” is a statement about deferred work, not
about work that never happens.
Famous related terms
- TLB —
TLB ≈ CPU cache for virtual→physical translations— the reason page-table walks aren’t crippling on every load. - Page fault —
page fault = MMU saying "this PTE isn't usable" + kernel handler that decides what to do— the universal hook for paging features. mmap—mmap = map a file (or anonymous pages) into your address space + let the kernel page it in on demand— virtual memory exposed as an API.- Copy-on-write —
CoW = share pages read-only + copy lazily on first write— what makesforkcheap. - Swap / paging —
swap ≈ disk used as overflow RAM, one page at a time— the original “more memory than the box has” trick. - OOM killer — see OOM killer — what happens when overcommit’s bet finally loses.
- Huge pages —
huge pages = 2 MB or 1 GB pages instead of 4 KB— fewer TLB entries cover more memory; helps big-RAM workloads.
Going deeper
- Linux kernel source:
mm/memory.candarch/x86/mm/fault.c— the primary source for “what does the kernel actually do when a fault arrives,” i.e. the real decision tree behind cases 1–5 above. - Operating Systems: Three Easy Pieces, the virtualization-of-memory chapters — go here if you want the whole model built up properly, with the paging math this post skipped; it’s free online.
- Ulrich Drepper, What Every Programmer Should Know About Memory — the rabbit hole: how much of your program’s speed is actually decided by TLBs and caches, i.e. where this abstraction leaks into performance.