Why fork() is such a weird API
Other systems take a program and arguments. Unix takes your whole process and clones it. The reasons are half historical accident, half deep insight — and the seams still show.
On this page
The picture version
Six pictures for a reader who has never called a system call, following one command you have typed a thousand times: ls | grep foo.
1 · The strange part
Your shell doesn’t start two programs. It copies itself twice.
2 · One call, two answers
The same line of code runs in two processes and takes different branches.
3 · Why copying is affordable
Nothing is copied until somebody writes.
4 · The gap is the point
Between the copy and the new program, you can just run code.
5 · Where it bites
The threads don’t come along. Their locks do.
6 · Keep this card
The whole thing on one index card.
Why it exists
Type ls | grep foo into a terminal and hit enter. What your shell does next
is not “start ls, start grep.” It makes two copies of itself — two
complete duplicate shells — and each copy then abandons its own program and
becomes ls or grep. That’s the running example for this post: a command you
have typed a thousand times, implemented by cloning the thing that’s reading
your keystrokes.
The first time you actually look at the call that does this, fork(), it
should feel wrong. You call one function, and it returns twice — once in the
original process, once in a new copy of the original process that magically
picks up at the exact same line of code. They’re separate processes with their
own memory from that point on — though on any modern system they start out
sharing the physical pages behind it, and their file descriptors refer to the
same open files, which is the detail the whole rest of this post rests on.
Nearly every other mainstream process-creation API models “make a new process” as
“give me a program and some arguments, and I’ll start it for you” — Windows
has CreateProcess, classic Mac OS / VMS / older mainframes had similar
spawn-style calls. Unix said: don’t pass me a program. Pass me yourself.
Choosing the program is a separate step afterwards, called exec, which
swaps a new program into the process you already have.
Why? The honest answer is that fork existed before there was a sensible alternative. The earliest Unix ran on a PDP-7 with no MMU and tiny memory. Dennis Ritchie’s own retrospective, The Evolution of the Unix Time-sharing System, is blunt about it: the system already swapped processes to disk, so fork needed little more than a bigger process table and a call that copied the current process into the swap area — “the PDP-7’s fork call required precisely 27 lines of assembly code.” He is careful to keep his own explanation a supposition rather than a verdict — it “seems reasonable to suppose that it exists in Unix mainly because of the ease with which fork could be implemented without changing much else” — but that is the closest thing to a designer’s account we have: a quick implementation choice, and then hard to dislodge once the rest of the system grew up around it.
But once it was there, people noticed it had a strange property: it cleanly separates creating a new process from deciding what that process will run. That separation is the deep insight, and it’s why fork outlived the PDP-7.
Why it matters now
Even if you never write fork() by hand, you live downstream of it:
- Every shell pipeline is fork-then-exec, repeated.
ls | grep foo | wcis three forks and three execs, with file descriptors wired between them in the gap between fork and exec. - Servers like nginx and PostgreSQL run pools of forked worker processes. Forking is what makes a worker cheap to start and what lets it inherit the parent’s setup — its configuration and its already-open listening sockets — without any of it being passed as arguments.
- Python’s GIL
workaround is process-based, and fork is the original engine. Reaching for
multiprocessingto use multiple cores means starting real OS processes — historically by forking. CPython’s default start method has been moving away from plain fork, precisely because of the hazards further down this post:spawnon macOS, andforkserveron other POSIX platforms as of Python 3.14. Check what your version does rather than trusting a remembered default. - Container runtimes are fork-with-extra-flags at the core. Linux’s
clone()is the generalized version, and namespaces — the thing that makes a container a container — are arguments to it. (Starting a container is more than that one call, but that call is the moment the container becomes a separate world.) - Fork’s weirdness is also why some things don’t work in
AI/ML stacks. PyTorch’s DataLoader can deadlock under the
forkstart method because the child inherits a CUDA context, threadpool, or malloc lock from the parent that’s now in an inconsistent state. The recommended fix isspawn(orforkserver) — which is, basically, “stop using fork.” Note that the start method you get comes from Python’s default for your platform and version, not from PyTorch, which is why the same code can behave differently on two machines.
The short answer
fork = duplicate the current process into two + give each a different return value
Picture to keep: not a factory that builds a new machine from a blueprint, but a photocopier that copies the machine mid-job — same half-finished page, same open drawers — and then hands one copy a note saying “you’re the copy.” Where the picture breaks: nothing is physically duplicated at first; the two processes share the same pages until one of them writes.
The parent gets back the child’s PID;
the child gets back zero. Same code, same memory contents, same open files
— from that instant on, two independent processes running the same program.
What you do after fork (usually exec to replace your program, or just
keep running) is the creative part.
How it works
The two-return-values trick
It’s not really two returns. It’s one syscall, two processes. The kernel
duplicates the calling process, schedules both, and arranges that when each
one resumes from the syscall, the return value register holds something
different — 0 in the child, the child’s PID in the parent. So this idiom:
pid_t pid = fork();
if (pid == 0) {
// child: replace ourselves with a new program
execvp("ls", argv);
} else if (pid > 0) {
// parent: wait for child to finish
waitpid(pid, &status, 0);
} else {
// fork failed
}
…is one piece of source code that compiles to one binary, but at runtime the
if branches differently in each process because the syscall handed them
different return values. Once you see it that way, it stops being weird:
fork returns once per process, not once per call.
Copy-on-write makes it cheap
The naive read of fork — “duplicate the entire process’s memory” — would be ruinously expensive for a process holding gigabytes of state. It isn’t, because of copy-on-write:
- After fork, parent and child share the same physical pages.
- The kernel marks every page read-only in both processes’ page tables.
- The first time either process writes to a page, the MMU traps, the kernel allocates a fresh physical page, copies the contents, and lets that process continue. Only the touched pages get duplicated.
So fork is roughly O(page-table size), not O(memory size). A server holding 16 GB of cache can fork a worker while copying only the page tables and whatever handful of pages the child then dirties. That is the whole reason a design shaped by the memory budget of a 1970s minicomputer is still viable on a machine with gigabytes of heap.
The fork+exec separation, and why it’s actually useful
The deep value of fork isn’t speed; it’s the gap between fork and exec. Between those two calls, the child is a fully-constructed process that hasn’t yet committed to a program. You can:
- Redirect file descriptors.
dup2(pipe_fd, STDOUT_FILENO)in the child, before exec, is how shells wire up pipelines. - Drop privileges.
setuid()to a less privileged user before exec. - Change directory, set environment variables, set resource limits, join a different process group.
A spawn-style API has to accept all of these as options up front — and Windows’
CreateProcess takes 10 parameters, several of them structs. (That reading of
the parameter list — it is what “configure the child without a gap” costs you —
is an interpretation of the API’s shape, not a claim about anyone’s stated
intent.) Fork-then-exec lets you configure the child by
running ordinary code in it. That’s elegant, and it’s why the gap is still
worth having even though posix_spawn offers a fused, more constrained
alternative.
Show the seams
- Fork in a multithreaded process is treacherous. The child only
inherits the calling thread. Every other thread vanishes — but their
locks remain held in the child’s memory, so the child can hang the moment it
touches
malloc(which itself takes locks). This is why POSIX restricts what you may legally do between fork and exec to async-signal-safe functions only. Many Python and PyTorch fork-related bugs trace back to this. - The “fork is cheap” claim has a footnote: page-table size. Forking a process with very large memory (terabytes, common in databases or ML training) can take noticeable wall-clock time just to copy the page tables, even with COW. Note what that cost scales with: how much memory is mapped, not how much the child goes on to touch.
vforkis the historical optimization for “I’m going to exec immediately anyway, don’t bother setting up COW.” It’s still in POSIX but is widely considered dangerous;posix_spawnis the modern answer.- macOS technically forks but discourages it. Apple’s documentation is
explicit that a process using its higher-level frameworks (Core Foundation,
Cocoa, Core Data) must call
execafter forking — the frameworks aren’t supported in a forked child that keeps running. The platform expectation is fork-then-exec orposix_spawn. - Containers complicate the story.
clone()(Linux’s superset of fork) takes flags that say “and also give the child a new mount namespace, new network namespace, new PID namespace…” That call is the heart of howruncstarts a container — fork with extra arguments — though there’s real setup around it.
So go back to ls | grep foo. Your shell forks twice; in each child, before
exec, it runs a few ordinary lines of C to point stdout at a pipe and stdin at
the other end; then each child execs its program and forgets it was ever a
shell. The pipeline isn’t a feature of ls or of grep, and the kernel
supplies only the raw pipe — the wiring that turns two unrelated programs into
one command is what the shell does in the gap.
You started with fork = duplicate the current process into two + give each a different return value. What did the pipeline add? — + the gap before exec.
Duplication is the historical accident; the gap is the design. Fork is weird
because it’s a clone primitive in a world that mostly builds spawn
primitives — and the reason it outlived its excuse is that a clone gives you a
fully-formed process you can configure by running code, instead of a struct with
ten fields.
Check yourself
Before you go — a web server forks 8 worker processes at startup, each inheriting a 4 GB in-memory cache the parent loaded. Someone reports total memory usage climbing steadily over the following hour even though the cache never grows. What’s your first hypothesis?
Answer
Copy-on-write is being undone one page at a time. The 8 workers started sharing
the parent’s 4 GB, so the initial cost really was near zero — but every page a
worker writes gets privately copied. And writes here don’t have to be your
data: a reference-counted runtime (CPython is the classic case) touches object
headers just by reading objects, which dirties the page and forces the copy.
Same for a garbage collector that marks in place. The fix is usually to keep the
shared data out of the runtime’s reach — mmaped read-only, or in a
representation the runtime doesn’t touch — not to fork fewer workers. (CPython
is the well-documented case; whether another runtime’s collector dirties shared
pages depends on how it marks, which varies by runtime.)
And one more — Python and PyTorch both warn against the fork start method in
programs that use threads. Given what fork copies,
why would a forked child hang on a lock nobody is holding? (The exact defaults
move between versions — the mechanism below is the durable part.)
Answer
Because the child inherits the parent’s memory, including the state of every
lock — but only one thread, the one that called fork. If another thread held
the malloc arena lock (or a logging lock, or a CUDA driver lock) at the instant
of the fork, the child gets a snapshot where that lock is held by a thread that
doesn’t exist in the child and can never release it. The next allocation in the
child blocks forever. Nobody is holding the lock in any meaningful sense; the
bits say it’s held. That’s the reason POSIX limits you to async-signal-safe
calls between fork and exec, and the reason spawn — which starts a fresh
process instead of cloning a snapshot — is the recommended escape.
Famous related terms
exec—exec = replace this process's program + keep PID, fds, and other process attributes— the other half of the fork+exec pair.clone()—clone = fork + flags for which resources to share or namespace— Linux’s generalized fork; threads and containers are both clone calls with different flags.posix_spawn—posix_spawn ≈ fork + exec fused into one call + a file-actions list— the modern, portable alternative when you don’t need the gap; it’s a library interface, so what it costs underneath varies by platform.- Copy-on-write —
COW = share pages read-only + allocate on first write— the trick that makes fork cheap; also used by filesystems like ZFS and btrfs. vfork—vfork ≈ fork that shares the parent's address space until exec— historical speed hack, easy to misuse.- Process —
process = address space + file descriptors + scheduling state + PID— the unit fork duplicates.
Going deeper
- Dennis Ritchie, The Evolution of the Unix Time-sharing System — the primary source, and the place to go for “was fork designed or improvised?”, answered by the person who was there.
- Baumann, Appavoo, Krieger, Roscoe, A fork() in the road (HotOS 2019) — the explainer to read if you want the full case against fork: a systematic inventory of every weirdness it introduces, argued by people who want it retired.
- The Linux
clone(2)man page — the rabbit hole: what fork looks like once it’s been generalized into “share exactly these resources and namespace those,” which is where threads and containers both come from.