Why memory-mapped files exist
Why hand a file to the OS as memory instead of reading it byte by byte? Because the OS was already caching it that way — and pretending otherwise costs you a copy you don't need. What you pay for deleting the copy is knowing when the I/O happens.
On this page
The picture version
Five pictures for a reader who has never wondered where a file lives while a program is reading it. The prose below fills in the seams the pictures skip.
1 · The problem
The bytes end up in memory twice.
2 · The move
Don’t ask for the bytes. Ask to see the room.
3 · What happens when you touch it
The array access is where the disk read hides.
4 · The bill
You bought all of this by giving up the schedule.
5 · Keep this card
The whole thing on one index card.
Why it exists
You download a 30 GB model file and start llama.cpp on it — on a laptop
with 16 GB of RAM, which should not be enough. It starts generating
anyway. Then you start a second copy against the same file and your memory
usage doesn’t double. Whatever happened there, it was not “read 30 GB of
weights into this process.”
You probably assume mmap means “load the file into RAM, but faster.” It’s
closer to the opposite. mmap doesn’t load the file at all; it arranges for
the bytes to arrive one page at a time, at the moment you touch them, into
memory the kernel already owns and can share.
Here’s the setup it’s reacting against. The obvious way to read that
weights file is open, then read(fd, buf, n) in a loop: you ask the OS
for some bytes, it puts them in your buffer, you use them. But your read
doesn’t usually reach the disk. The kernel keeps a
page cache —
an in-memory cache of file pages maintained for everyone — and once a page
is in it, read is a copy out of that cache into your buffer. (On a
cold cache the kernel does go to disk first, then copies. Either way you
end the call with two copies of those bytes in physical memory.) You spent
a syscall
plus a memcpy for the privilege. For a config file this is irrelevant. For
30 GB of weights that you’ll read large chunks of, it’s the whole ballgame.
mmap is what you reach for when you decide the copy is the problem.
Instead of
“give me the bytes,” you tell the kernel: map this file into my address
space. From then on, the weights file is just a pointer. Reading byte
1,000,000 of it is addr[1000000]. The page cache pages and your “buffer”
are literally the same physical pages — the kernel has just made them appear
in your virtual address space too.
That’s the deep idea: the page cache is already a memory-mapped view of the
file from the kernel’s side. mmap just lets you skip the pretense that
your process is somewhere else.
Why it matters now
The clearest modern case for mmap is a file too big to comfortably copy:
- Loading model weights.
llama.cpploads GGUF weights throughmmapwherever the backend it is running on supports doing so, which is the default mode. The 30 GB file isn’t loaded into RAM up front; the OS pulls in pages on demand as the inference engine touches tensor data. Multiple processes loading the same model can share the same physical pages because the page cache is shared — the first process pays the disk read; the rest hit warm cache, while those pages stay resident. safetensorsoffers a memory-mapped loading mode for the same reason — zero-copy reads, lazy faulting, no doubled RAM.- Databases. LMDB is built around
mmap. SQLite supports it as an opt-in mode (off by default). MongoDB’s old MMAPv1 engine was named after it. Postgres deliberately goes the other way and manages its own buffer pool, for reasons the Check yourself section below comes back to. There are good arguments on both sides. - Executables. The OS loader maps ELF
binaries and shared libraries into
memory rather than reading them; that’s why two processes running the
same binary share its read-only pages. Arrow IPC and other random-access
binary formats are also commonly consumed via
mmap. - Inter-process communication via shared memory is a special case of
the same primitive:
mmapsomething, multiple processes see the same bytes.
If your job involves loading or scanning files bigger than a few megabytes,
the version of you with mmap in their toolkit makes meaningfully
different design choices than the one without.
The short answer
mmap = the kernel exposes the page cache as part of your virtual address space + pages get loaded on demand via page faults
Picture to keep: not a truck delivering the file to your house — a window cut into the room where the kernel keeps it. Two processes can cut windows into the same room and see the same furniture. Where the picture breaks: a window shows you what’s already there, whereas touching an unmapped page makes the kernel go fetch it from disk — the fetch is real I/O, it just happens inside what looks like an array access.
You stop calling read. The file just is a region of your virtual memory.
Touching a byte that hasn’t been loaded yet triggers a
page fault,
the kernel pulls the relevant page in from disk into the page cache, and
your access continues. Pages you never touch never get loaded. Pages other
processes are using are shared. The copy goes away because there was never
a separate buffer to copy into.
How it works
Attempt 1: read the weights into a buffer
malloc(30 GB), then read the file into it. This fails three ways at
once on the 16 GB laptop: you need 30 GB of anonymous memory you don’t
have, you pay the copy out of the page cache, and the second llama.cpp
process needs its own 30 GB because a private buffer is private. You also
wait for the whole file before answering the first token, even though
inference touches tensors in a pattern that leaves plenty of the file
untouched for a while.
Every one of those is a consequence of the same decision: the bytes live in a buffer that belongs to your process. So don’t have a buffer.
Attempt 2: map it instead
mmap(addr, len, prot, flags, fd, offset) returns a pointer. The kernel
sets up page table
entries that say: “these virtual addresses correspond to that range of
that file.” Crucially, it does not read the file yet (unless you pass
MAP_POPULATE, which asks the kernel to populate the page tables and read
ahead up front). The page table entries are marked not-present.
Everything below is Linux-specific in its details — flag names, madvise
semantics, what counts as a fault. The shape — map, fault, share — is
common to other Unixes and to Windows’ file-mapping API, but don’t port the
specifics without checking.
The first time your code dereferences one of those addresses, the MMU sees the not-present entry and raises a page fault. The kernel’s page-fault handler:
- Looks up which file and offset this virtual address corresponds to.
- Checks the page cache. If the page is already there (because someone else read this file recently), great — no disk read at all.
- Otherwise, issues a disk read for that page (typically 4 KB, possibly more if readahead kicks in).
- Updates your page table entry to point at the page-cache page.
- Returns from the fault. Your instruction retries and now succeeds.
The same physical 4 KB page in RAM is now reachable from the kernel’s page cache, from your virtual address space, and from any other process that has mapped the same file. For an ordinary file mapping on Linux, that’s one copy of the data instead of two.
For MAP_SHARED mappings, writes go back to the file lazily, via the
kernel’s normal dirty-page writeback — when they’re actually durable on
disk is a separate question (msync / fsync). For MAP_PRIVATE, writes are
copy-on-write —
your dirty pages get a private copy, the rest stay shared.
What that bought
The three failures from attempt 1, in order:
- No copy.
readdoes kernel-page-cache → user-buffer.mmapskips the destination buffer entirely. - Lazy loading. The 30 GB file mapped in is essentially free until you
touch it — which is why
llama.cppanswers before it has read the whole thing. The kernel pulls in only what you access. - Sharing across processes. Two processes that map the same file see the same physical pages. Starting the second inference worker on the same weights doesn’t double your RAM bill.
- The OS handles eviction for you. Under memory pressure, the kernel can drop clean mapped pages without consulting you; they’ll fault back in if needed. You get something close to “the file is in RAM when there’s RAM, on disk when there isn’t” without writing any caching code.
What it broke on the way
Every one of those wins came from handing control of the I/O to the kernel. That’s also the bill.
- I/O errors become signals, not return values.
readreturns-1witherrnoset. A fault inside anmmapregion — accessing past the end of a file that got truncated under you is the canonical case — can surface as SIGBUS from inside an ordinary memory access. Code that wasn’t expecting “loading this byte might fail” handles this badly. This is one of the real reasons some serious databases avoidmmap. - You don’t control when I/O happens. With
readyou decide. Withmmap, every memory access is a possible disk read. A latency-sensitive thread can stall on a page fault you didn’t see coming. - Random access patterns are TLB-hostile. Each 4 KB page you touch
may consume a TLB
entry. Sweeping a multi-GB file across a small TLB causes a lot of TLB
misses and the page-walk cost that comes with them. Huge pages attack
this directly — one entry covers 2 MB instead of 4 KB.
madvisedoesn’t help the TLB; it helps the other problem, which is the kernel guessing wrong about what to read ahead. - Address-space exhaustion was a real thing on 32-bit. A 4 GB virtual
address space couldn’t map a 5 GB file. This is one of the reasons
mmap-based storage engines were rough on 32-bit systems; on 64-bit it’s a non-issue for any realistic file. - The kernel may make worse caching decisions than your app. Postgres’
case for managing its own buffer pool is exactly this: a database knows
which pages are hot in a way the kernel’s generic page-replacement
policy does not.
mmapis great when “treat the OS’s page cache as your cache” is right; less great when it isn’t. madviseis the steering wheel.MADV_SEQUENTIALtells the kernel to read ahead aggressively and drop pages quickly;MADV_RANDOMsuppresses readahead;MADV_DONTNEEDasks the kernel to drop the pages (semantics differ for shared vs. anonymous mappings). If yourmmapworkload feels slow, a missingmadvisehint is one of the first things to check.
You started with mmap = the kernel exposes the page cache as part of your virtual address space + pages get loaded on demand. Before you scroll —
what did the 30 GB weights file add to that line, that isn’t in it yet?
— + you gave up knowing when I/O happens. That single trade explains both halves: it’s why the second worker is free
and why a SIGBUS can come out of an ordinary array read. When mmap
wins it’s not because it’s clever, it’s because it deleted a step. When it
loses it’s because your control flow is now welded to the kernel’s
page-fault and writeback behaviour. Pick the one whose failure modes you’d
rather debug.
Check yourself
Before you go — you mmap the 30 GB weights file, and inference is fast.
Your colleague mmaps the same file on a box where /models is an
NFS
mount over a busy network, and their p99 token latency is terrible even
though the file is identical. What changed?
Answer
Nothing about mmap changed — what changed is the cost of a page fault.
Every untouched page is a synchronous fetch at the moment your code
dereferences it, and on a network filesystem that fetch is slow and jittery
in a way a local NVMe read isn’t. With read you’d at least have known
where the I/O was, and could have prefetched or batched it; with mmap
the stall is hidden inside a memory access on the inference thread. This is
the general shape of “mmap is unpredictable”: the win and the risk are the
same mechanism.
And a trade-off: if mmap avoids a copy and lets the kernel handle
eviction, why does Postgres deliberately manage its own buffer pool instead?
Answer
Because “let the kernel decide what stays in RAM” is only a win when the
kernel’s page-replacement policy is as good as yours. The kernel’s policy
is general-purpose and has to serve every process on the box; a database
knows things it can’t — which pages are index roots, which scan is a
one-off that shouldn’t evict the working set — and it needs to control
exactly when a dirty page reaches disk, because
write-ahead-log
ordering depends on it. Postgres trades the saved copy for control over caching and
durability. Note the shape of the trade: it isn’t that mmap is slow, it’s
that it takes decisions away from a program that had better information.
Famous related terms
- Page cache —
page cache = kernel-managed RAM cache of file pages, keyed by (file, offset)— the thingmmapexposes; without it,mmapwould have nothing to hand you. - Demand paging —
demand paging = pages get loaded only when first accessed, via page fault— the load-on-touch behavior that makes mapping huge files cheap. - Copy-on-write —
COW = shared page until someone writes; then duplicate lazily— whatMAP_PRIVATEandforkboth rely on. - Huge pages —
huge pages = 2 MB or 1 GB page-table entries instead of 4 KB— reduce TLB pressure when mapping large regions; relevant for big model files. madvise—madvise = hints to the kernel about your access pattern—SEQUENTIAL,RANDOM,WILLNEED,DONTNEED. The way you tune anmmapworkload.- Zero-copy I/O —
zero-copy = avoid the kernel↔user memcpy on the I/O path—mmapis the page-cache-sharing flavor;sendfile,splice, andio_uringregistered buffers attack the same overhead from different angles. - Virtual memory — the substrate that makes any of this possible.
- Syscalls are expensive — part of
mmap’s win is one syscall up front instead of one perread.
Going deeper
- The Linux
mmap(2)andmadvise(2)man pages — the primary source for what the flags actually promise; most of the gotchas above are quietly documented there, including theSIGBUScase. - “Are You Sure You Want to Use MMAP in Your Database Management System?”
(Crotty, Leis, Pavlo, CIDR 2022) — the best end-to-end explanation of
“when is
mmapthe wrong tool,” written by people who wanted it to work and enumerate exactly where it didn’t. - The
llama.cppsource (search formmap) — the rabbit hole, if you want to see what the trade-offs above look like in a real consumer that bet on this primitive.