Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why fork() is such a weird API

Other systems take a program and arguments. Unix takes your whole process and clones it. The reasons are half historical accident, half deep insight — and the seams still show.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 13 min read

On this page

The picture version

Six pictures for a reader who has never called a system call, following one command you have typed a thousand times: ls | grep foo.

1 · The strange part

Your shell doesn’t start two programs. It copies itself twice.

your shell reading your keystrokes copies itself a duplicate shell a duplicate shell becomes becomes ls grep Making a new process and choosing its program are two separate steps. every other mainstream API fuses them: hand me a program and some arguments, and I will start it for you
Unix went the other way: don’t pass me a program, pass me yourself. The duplicate is a real, running process that has not yet decided what it will be — and that undecided moment turns out to be the useful part.

2 · One call, two answers

The same line of code runs in two processes and takes different branches.

pid = fork(); in the parent pid = the child’s id in the child pid = 0 It is not one call returning twice. It is one call, in two processes. the kernel duplicates the caller and hands each copy a different answer, so a single if sorts them out
One source file, one binary, two processes running it. fork returns once per process, not once per call — and the return value is the only thing telling a process which one it is.

3 · Why copying is affordable

Nothing is copied until somebody writes.

parent child the same physical pages marked read-only in both then one of them writes to a page the hardware traps, because the page was marked read-only the kernel makes a private copy of that page only and lets the write through every page nobody writes is never duplicated at all cost ≈ O(page tables), not O(memory) a server holding 16 GB of cache forks a worker by copying page tables and whatever handful of pages the child dirties with a footnote: the cost tracks how much memory is mapped, so a terabyte-sized process still pays to copy the tables
Copy-on-write is what keeps a 1970s design viable on a machine with gigabytes of heap. “Duplicate the process” is a promise about what the child sees, not about what was actually copied.

4 · The gap is the point

Between the copy and the new program, you can just run code.

fork the gap exec point stdout at a pipe — this is the entire trick behind ls | grep foo drop privileges to a less powerful user change directory · set environment variables · set resource limits A spawn-style API has to take all of that as parameters up front. Windows’ CreateProcess takes ten of them, several being structs fork configures the child by running ordinary code inside it
Duplication is the historical accident; the gap is the design. The pipeline is not a feature of ls or of grep — the kernel supplies a raw pipe, and the wiring happens in those few lines before exec.

5 · Where it bites

The threads don’t come along. Their locks do.

the parent, mid-fork thread 1 — calls fork thread 2 — holding the malloc lock thread 3 — holding a logging lock perfectly healthy the child, one instant later thread 1 — the only one that came along the malloc lock: still marked held by a thread that does not exist here so the next allocation blocks forever Nobody is holding the lock. The bits say it is held. memory came across whole; threads did not — which is why POSIX limits you to async-signal-safe calls before exec it is also why Python and PyTorch push you toward spawn and forkserver: a fresh process instead of a snapshot
The seam is exactly the property that made fork elegant. Cloning a process means cloning a moment, and a moment that was consistent for five threads is not consistent for the one that survives.

6 · Keep this card

The whole thing on one index card.

fork = duplicate the current process into two + give each a different return value + the gap before exec duplication is the accident; the gap is the design a clone primitive in a world of spawn primitives — and it survived because a clone is a process you configure by running code in it
Picture to keep: not a factory building a new machine from a blueprint, but a photocopier that copies the machine mid-job — same half-finished page, same open drawers — and then hands one copy a note saying “you’re the copy.” Where it breaks: nothing is physically duplicated at first; the two share the same pages until one writes.

Why it exists

Type ls | grep foo into a terminal and hit enter. What your shell does next is not “start ls, start grep.” It makes two copies of itself — two complete duplicate shells — and each copy then abandons its own program and becomes ls or grep. That’s the running example for this post: a command you have typed a thousand times, implemented by cloning the thing that’s reading your keystrokes.

The first time you actually look at the call that does this, fork(), it should feel wrong. You call one function, and it returns twice — once in the original process, once in a new copy of the original process that magically picks up at the exact same line of code. They’re separate processes with their own memory from that point on — though on any modern system they start out sharing the physical pages behind it, and their file descriptors refer to the same open files, which is the detail the whole rest of this post rests on.

Nearly every other mainstream process-creation API models “make a new process” as “give me a program and some arguments, and I’ll start it for you” — Windows has CreateProcess, classic Mac OS / VMS / older mainframes had similar spawn-style calls. Unix said: don’t pass me a program. Pass me yourself. Choosing the program is a separate step afterwards, called exec, which swaps a new program into the process you already have.

Why? The honest answer is that fork existed before there was a sensible alternative. The earliest Unix ran on a PDP-7 with no MMU and tiny memory. Dennis Ritchie’s own retrospective, The Evolution of the Unix Time-sharing System, is blunt about it: the system already swapped processes to disk, so fork needed little more than a bigger process table and a call that copied the current process into the swap area — “the PDP-7’s fork call required precisely 27 lines of assembly code.” He is careful to keep his own explanation a supposition rather than a verdict — it “seems reasonable to suppose that it exists in Unix mainly because of the ease with which fork could be implemented without changing much else” — but that is the closest thing to a designer’s account we have: a quick implementation choice, and then hard to dislodge once the rest of the system grew up around it.

But once it was there, people noticed it had a strange property: it cleanly separates creating a new process from deciding what that process will run. That separation is the deep insight, and it’s why fork outlived the PDP-7.

Why it matters now

Even if you never write fork() by hand, you live downstream of it:

The short answer

fork = duplicate the current process into two + give each a different return value

Picture to keep: not a factory that builds a new machine from a blueprint, but a photocopier that copies the machine mid-job — same half-finished page, same open drawers — and then hands one copy a note saying “you’re the copy.” Where the picture breaks: nothing is physically duplicated at first; the two processes share the same pages until one of them writes.

The parent gets back the child’s PID; the child gets back zero. Same code, same memory contents, same open files — from that instant on, two independent processes running the same program. What you do after fork (usually exec to replace your program, or just keep running) is the creative part.

How it works

The two-return-values trick

It’s not really two returns. It’s one syscall, two processes. The kernel duplicates the calling process, schedules both, and arranges that when each one resumes from the syscall, the return value register holds something different — 0 in the child, the child’s PID in the parent. So this idiom:

pid_t pid = fork();
if (pid == 0) {
    // child: replace ourselves with a new program
    execvp("ls", argv);
} else if (pid > 0) {
    // parent: wait for child to finish
    waitpid(pid, &status, 0);
} else {
    // fork failed
}

…is one piece of source code that compiles to one binary, but at runtime the if branches differently in each process because the syscall handed them different return values. Once you see it that way, it stops being weird: fork returns once per process, not once per call.

Copy-on-write makes it cheap

The naive read of fork — “duplicate the entire process’s memory” — would be ruinously expensive for a process holding gigabytes of state. It isn’t, because of copy-on-write:

  1. After fork, parent and child share the same physical pages.
  2. The kernel marks every page read-only in both processes’ page tables.
  3. The first time either process writes to a page, the MMU traps, the kernel allocates a fresh physical page, copies the contents, and lets that process continue. Only the touched pages get duplicated.

So fork is roughly O(page-table size), not O(memory size). A server holding 16 GB of cache can fork a worker while copying only the page tables and whatever handful of pages the child then dirties. That is the whole reason a design shaped by the memory budget of a 1970s minicomputer is still viable on a machine with gigabytes of heap.

The fork+exec separation, and why it’s actually useful

The deep value of fork isn’t speed; it’s the gap between fork and exec. Between those two calls, the child is a fully-constructed process that hasn’t yet committed to a program. You can:

A spawn-style API has to accept all of these as options up front — and Windows’ CreateProcess takes 10 parameters, several of them structs. (That reading of the parameter list — it is what “configure the child without a gap” costs you — is an interpretation of the API’s shape, not a claim about anyone’s stated intent.) Fork-then-exec lets you configure the child by running ordinary code in it. That’s elegant, and it’s why the gap is still worth having even though posix_spawn offers a fused, more constrained alternative.

Show the seams

So go back to ls | grep foo. Your shell forks twice; in each child, before exec, it runs a few ordinary lines of C to point stdout at a pipe and stdin at the other end; then each child execs its program and forgets it was ever a shell. The pipeline isn’t a feature of ls or of grep, and the kernel supplies only the raw pipe — the wiring that turns two unrelated programs into one command is what the shell does in the gap.

You started with fork = duplicate the current process into two + give each a different return value. What did the pipeline add? — + the gap before exec. Duplication is the historical accident; the gap is the design. Fork is weird because it’s a clone primitive in a world that mostly builds spawn primitives — and the reason it outlived its excuse is that a clone gives you a fully-formed process you can configure by running code, instead of a struct with ten fields.

Check yourself

Before you go — a web server forks 8 worker processes at startup, each inheriting a 4 GB in-memory cache the parent loaded. Someone reports total memory usage climbing steadily over the following hour even though the cache never grows. What’s your first hypothesis?

Answer

Copy-on-write is being undone one page at a time. The 8 workers started sharing the parent’s 4 GB, so the initial cost really was near zero — but every page a worker writes gets privately copied. And writes here don’t have to be your data: a reference-counted runtime (CPython is the classic case) touches object headers just by reading objects, which dirties the page and forces the copy. Same for a garbage collector that marks in place. The fix is usually to keep the shared data out of the runtime’s reach — mmaped read-only, or in a representation the runtime doesn’t touch — not to fork fewer workers. (CPython is the well-documented case; whether another runtime’s collector dirties shared pages depends on how it marks, which varies by runtime.)

And one more — Python and PyTorch both warn against the fork start method in programs that use threads. Given what fork copies, why would a forked child hang on a lock nobody is holding? (The exact defaults move between versions — the mechanism below is the durable part.)

Answer

Because the child inherits the parent’s memory, including the state of every lock — but only one thread, the one that called fork. If another thread held the malloc arena lock (or a logging lock, or a CUDA driver lock) at the instant of the fork, the child gets a snapshot where that lock is held by a thread that doesn’t exist in the child and can never release it. The next allocation in the child blocks forever. Nobody is holding the lock in any meaningful sense; the bits say it’s held. That’s the reason POSIX limits you to async-signal-safe calls between fork and exec, and the reason spawn — which starts a fresh process instead of cloning a snapshot — is the recommended escape.

Going deeper