Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why memory-mapped files exist

Why hand a file to the OS as memory instead of reading it byte by byte? Because the OS was already caching it that way — and pretending otherwise costs you a copy you don't need. What you pay for deleting the copy is knowing when the I/O happens.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Five pictures for a reader who has never wondered where a file lives while a program is reading it. The prose below fills in the seams the pictures skip.

1 · The problem

The bytes end up in memory twice.

the file on disk, 30 GB read once the kernel’s page cache RAM the kernel keeps for everyone copied your buffer the same bytes, again every read() ends with two copies of those bytes in memory for a config file, who cares for 30 GB of model weights on a 16 GB laptop, the copy is the entire problem and a second program reading the same file allocates its own 30 GB, because a private buffer is private The buffer belonging to your process is the thing causing all three failures.
Nothing here is wasteful by accident — a buffer you own is exactly what read() promises. The cost is that the kernel already had those bytes in RAM, and handing you a copy is the only way to keep that promise.

2 · The move

Don’t ask for the bytes. Ask to see the room.

one set of physical pages the kernel’s page cache, holding the file process A sees it at some address process B sees it at a different one no copy no copy the file isn’t loaded into your memory — your memory is pointed at the file picture a window cut into the room where the kernel keeps the file, not a truck delivering it to your house two processes can cut windows into the same room, and the furniture is not duplicated
This is why the second inference worker is nearly free, and why the 30 GB file fits on a 16 GB laptop at all. The page cache was always a memory-mapped view of the file from the kernel’s side — mapping just drops the pretence that your process lives somewhere else.

3 · What happens when you touch it

The array access is where the disk read hides.

addr[1000000] an ordinary read of memory nothing is there yet the CPU raises a page fault the kernel takes over 1  which file and offset does this address mean? 2  is that page already in the page cache? — if yes, no disk read happens at all 3  if not, read it from disk, usually with some readahead for its neighbours 4  point this process’s page table at that page, and let the instruction run again Pages you never touch are never read. Pages someone else touched are already warm. which is how generation starts before anything has read the whole 30 GB
Step 2 is the one worth keeping: the fault is a lookup first and a disk read only if the lookup misses. The first process to touch a page pays for it; every later reader of that same file gets it for the price of a page-table update.

4 · The bill

You bought all of this by giving up the schedule.

what it bought what it broke no copy — there is no second buffer to fill lazy — untouched pages cost nothing shared — a second process pays almost nothing the kernel evicts clean pages under pressure, and they fault back in if you need them again an I/O error arrives as a signal, mid-instruction, rather than as a value you can check any memory access can stall on a disk read you didn’t schedule and can’t see the kernel’s caching policy is general-purpose, and your program may know better than it does sweeping a huge file thrashes address translation Both columns are the same sentence: the kernel is now doing your I/O.
This is why the answer to “should I use mmap?” is never general. A database that knows which pages are hot is giving up real information by delegating, while an inference engine reading a huge immutable file is giving up nothing it wanted.

5 · Keep this card

The whole thing on one index card.

mmap = the kernel exposes the page cache inside your address space + pages get loaded on demand + you gave up knowing when I/O happens that one trade explains the free second worker and the crash that comes out of an ordinary array read when it wins, it isn’t being clever — it deleted a step when it loses, your control flow is welded to the kernel’s
So the choice isn’t fast versus slow. It is which set of failure modes you would rather debug — a copy you can see and schedule, or a stall and a signal you cannot.

Why it exists

You download a 30 GB model file and start llama.cpp on it — on a laptop with 16 GB of RAM, which should not be enough. It starts generating anyway. Then you start a second copy against the same file and your memory usage doesn’t double. Whatever happened there, it was not “read 30 GB of weights into this process.”

You probably assume mmap means “load the file into RAM, but faster.” It’s closer to the opposite. mmap doesn’t load the file at all; it arranges for the bytes to arrive one page at a time, at the moment you touch them, into memory the kernel already owns and can share.

Here’s the setup it’s reacting against. The obvious way to read that weights file is open, then read(fd, buf, n) in a loop: you ask the OS for some bytes, it puts them in your buffer, you use them. But your read doesn’t usually reach the disk. The kernel keeps a page cache — an in-memory cache of file pages maintained for everyone — and once a page is in it, read is a copy out of that cache into your buffer. (On a cold cache the kernel does go to disk first, then copies. Either way you end the call with two copies of those bytes in physical memory.) You spent a syscall plus a memcpy for the privilege. For a config file this is irrelevant. For 30 GB of weights that you’ll read large chunks of, it’s the whole ballgame.

mmap is what you reach for when you decide the copy is the problem. Instead of “give me the bytes,” you tell the kernel: map this file into my address space. From then on, the weights file is just a pointer. Reading byte 1,000,000 of it is addr[1000000]. The page cache pages and your “buffer” are literally the same physical pages — the kernel has just made them appear in your virtual address space too.

That’s the deep idea: the page cache is already a memory-mapped view of the file from the kernel’s side. mmap just lets you skip the pretense that your process is somewhere else.

Why it matters now

The clearest modern case for mmap is a file too big to comfortably copy:

If your job involves loading or scanning files bigger than a few megabytes, the version of you with mmap in their toolkit makes meaningfully different design choices than the one without.

The short answer

mmap = the kernel exposes the page cache as part of your virtual address space + pages get loaded on demand via page faults

Picture to keep: not a truck delivering the file to your house — a window cut into the room where the kernel keeps it. Two processes can cut windows into the same room and see the same furniture. Where the picture breaks: a window shows you what’s already there, whereas touching an unmapped page makes the kernel go fetch it from disk — the fetch is real I/O, it just happens inside what looks like an array access.

You stop calling read. The file just is a region of your virtual memory. Touching a byte that hasn’t been loaded yet triggers a page fault, the kernel pulls the relevant page in from disk into the page cache, and your access continues. Pages you never touch never get loaded. Pages other processes are using are shared. The copy goes away because there was never a separate buffer to copy into.

How it works

Attempt 1: read the weights into a buffer

malloc(30 GB), then read the file into it. This fails three ways at once on the 16 GB laptop: you need 30 GB of anonymous memory you don’t have, you pay the copy out of the page cache, and the second llama.cpp process needs its own 30 GB because a private buffer is private. You also wait for the whole file before answering the first token, even though inference touches tensors in a pattern that leaves plenty of the file untouched for a while.

Every one of those is a consequence of the same decision: the bytes live in a buffer that belongs to your process. So don’t have a buffer.

Attempt 2: map it instead

mmap(addr, len, prot, flags, fd, offset) returns a pointer. The kernel sets up page table entries that say: “these virtual addresses correspond to that range of that file.” Crucially, it does not read the file yet (unless you pass MAP_POPULATE, which asks the kernel to populate the page tables and read ahead up front). The page table entries are marked not-present.

Everything below is Linux-specific in its details — flag names, madvise semantics, what counts as a fault. The shape — map, fault, share — is common to other Unixes and to Windows’ file-mapping API, but don’t port the specifics without checking.

The first time your code dereferences one of those addresses, the MMU sees the not-present entry and raises a page fault. The kernel’s page-fault handler:

  1. Looks up which file and offset this virtual address corresponds to.
  2. Checks the page cache. If the page is already there (because someone else read this file recently), great — no disk read at all.
  3. Otherwise, issues a disk read for that page (typically 4 KB, possibly more if readahead kicks in).
  4. Updates your page table entry to point at the page-cache page.
  5. Returns from the fault. Your instruction retries and now succeeds.

The same physical 4 KB page in RAM is now reachable from the kernel’s page cache, from your virtual address space, and from any other process that has mapped the same file. For an ordinary file mapping on Linux, that’s one copy of the data instead of two.

For MAP_SHARED mappings, writes go back to the file lazily, via the kernel’s normal dirty-page writeback — when they’re actually durable on disk is a separate question (msync / fsync). For MAP_PRIVATE, writes are copy-on-write — your dirty pages get a private copy, the rest stay shared.

What that bought

The three failures from attempt 1, in order:

What it broke on the way

Every one of those wins came from handing control of the I/O to the kernel. That’s also the bill.

You started with mmap = the kernel exposes the page cache as part of your virtual address space + pages get loaded on demand. Before you scroll — what did the 30 GB weights file add to that line, that isn’t in it yet?

— + you gave up knowing when I/O happens. That single trade explains both halves: it’s why the second worker is free and why a SIGBUS can come out of an ordinary array read. When mmap wins it’s not because it’s clever, it’s because it deleted a step. When it loses it’s because your control flow is now welded to the kernel’s page-fault and writeback behaviour. Pick the one whose failure modes you’d rather debug.

Check yourself

Before you go — you mmap the 30 GB weights file, and inference is fast. Your colleague mmaps the same file on a box where /models is an NFS mount over a busy network, and their p99 token latency is terrible even though the file is identical. What changed?

Answer

Nothing about mmap changed — what changed is the cost of a page fault. Every untouched page is a synchronous fetch at the moment your code dereferences it, and on a network filesystem that fetch is slow and jittery in a way a local NVMe read isn’t. With read you’d at least have known where the I/O was, and could have prefetched or batched it; with mmap the stall is hidden inside a memory access on the inference thread. This is the general shape of “mmap is unpredictable”: the win and the risk are the same mechanism.

And a trade-off: if mmap avoids a copy and lets the kernel handle eviction, why does Postgres deliberately manage its own buffer pool instead?

Answer

Because “let the kernel decide what stays in RAM” is only a win when the kernel’s page-replacement policy is as good as yours. The kernel’s policy is general-purpose and has to serve every process on the box; a database knows things it can’t — which pages are index roots, which scan is a one-off that shouldn’t evict the working set — and it needs to control exactly when a dirty page reaches disk, because write-ahead-log ordering depends on it. Postgres trades the saved copy for control over caching and durability. Note the shape of the trade: it isn’t that mmap is slow, it’s that it takes decisions away from a program that had better information.

Going deeper