Why Linux has an OOM killer
Linux promises memory it doesn't have, then has to break the promise — the OOM killer is the reaper that decides who dies so the system can keep running.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Attempt 1: don’t overcommit — fail honestly
- Attempt 2: overcommit, and deal with it later
- Attempt 3: kill somebody — but who?
- Attempt 4: scope the whole thing to a cgroup
- Show the seams
- Check yourself
- Famous related terms
- Going deeper
The picture version
The whole idea in six pictures, for a reader who has never seen a process disappear without saying anything. The prose below fills in the seams.
1 · The problem
The log just stops. Then a number.
2 · The bet
Linux promises more memory than it has.
3 · When the bet loses
The bill arrives long after the sale.
4 · Choosing the victim
Score everybody. Kill the worst.
5 · The second kind of “out of memory”
The host is fine. Your container is not.
6 · Keep this card
The whole thing on one index card.
Why it exists
Your container is gone. No stack trace, no exception, no “out of memory” line in the application log — the last thing it printed was ordinary work, and then nothing. The orchestrator says the process exited with code 137. We’ll follow that container the whole way down.
You probably assume that running out of memory means malloc returns
NULL and your program handles it. On Linux, usually not. Your program
never found out it was out of memory; something else decided it should
stop existing.
Here’s why. A process calls malloc(8 GB) and the kernel says yes. Run a
few of these and total “allocated” memory exceeds physical RAM plus swap.
Nothing has crashed, because none of those processes have touched most of
that memory — malloc reserved address space, not pages.
Then one of them starts writing. The kernel has to find a real, physical
page to back each virtual page being touched. At some point it can’t: every
page is in use, swap is full, the page cache is already squeezed flat. The
kernel made a promise it can’t keep, and — this is the key part — it’s
discovering that during a memory access, not during malloc. There’s no
error code to return to anyone. The malloc that lied already succeeded,
possibly minutes ago.
Two options exist at that moment. One: panic the whole machine — every process dies, the box reboots. Two: pick a victim, kill it, free its pages, and let everyone else live. Linux chose option two and built the OOM killer to do it. Exit code 137 is your container being option two.
The OOM killer exists because Linux deliberately overcommits memory. Overcommit is a calculated bet that programs reserve far more than they touch — and most of the time the bet pays off. The OOM killer is what happens when it doesn’t.
Why it matters now
If you run anything in containers — Kubernetes pods, Docker containers, an
ML training job inside a memory-limited
cgroup
— you have already met the OOM killer. That exit code 137 decodes as
128 + 9: the shell convention for “terminated by signal 9,” and signal 9 is
SIGKILL, the one signal a process cannot catch, block, or ignore. Nothing in
your program was consulted. Worth knowing the boundary: 137 means something
sent SIGKILL — a kill -9, a runtime tearing the container down, an
orchestrator timing it out. OOM is the usual suspect in a container, not the
only one.
It matters even more now because of large models. An LLM that loads 70 GB of weights, a fine-tuning run that spikes activation memory, a vector index that grows past what you sized for — these are exactly the workloads that will push real pages into existence and force the kernel to make good on its promises. AI infra is OOM-killer infra.
The short answer
OOM killer = "system is out of memory" trigger + heuristic that scores processes + SIGKILL on the highest scorer
Picture to keep: an overbooked flight. The airline sold more seats than the plane has, betting some passengers won’t show. When they all show, the gate agent doesn’t cancel the flight — they pick someone and bump them. Where it breaks: the gate agent bumps by who’s cheapest to bump, whereas the kernel bumps by who’s biggest, which is often the one passenger the flight existed for.
When allocation fails and there’s nowhere left to reclaim from, the kernel
walks the process list, computes a “badness” score for each one, and sends
SIGKILL to the worst offender. The score roughly tracks how much memory the
process is using, with adjustments for privilege and policy.
How it works
Attempt 1: don’t overcommit — fail honestly
The kernel could refuse any malloc it couldn’t back with real RAM plus
swap. No lying, no reaper, malloc returns NULL and your program decides
what to do.
Why it breaks: two ways. Programs routinely reserve far more than they
touch — a fork of a 10 GB process, a sparse array, a thread stack — so
honest accounting refuses allocations that would have been fine and leaves
real RAM idle. And almost nobody checks malloc’s return value, so
“honest” failure mostly manifests as a null-pointer crash at a random
location instead of a clean error. Linux still offers this mode
(vm.overcommit_memory=2); most systems don’t run it.
Attempt 2: overcommit, and deal with it later
Overcommit takes the bet. Say yes, back pages only when touched, and reclaim aggressively — evict page cache, swap out cold anonymous pages — when things get tight.
Why it breaks: sometimes reclaim isn’t enough. The trigger is a failed page allocation under __alloc_pages that the kernel cannot satisfy even after reclaiming. Now it’s mid-page-fault with nothing to hand back and no caller who can cope. Something has to die.
Attempt 3: kill somebody — but who?
Killing at random works, technically. It also means a one-line log-shipper sidecar can be sacrificed while the 60 GB leaker that caused the crisis survives and triggers the next one thirty seconds later.
So selection uses a score per process, exposed at /proc/<pid>/oom_score. The
current formula dates from the badness-heuristic rewrite in Linux 2.6.36,
which is also when /proc/<pid>/oom_score_adj replaced the older oom_adj.
It is roughly:
- Base score = the memory the process actually occupies — resident pages, plus its swapped-out pages, plus the page tables describing them — as a fraction of the memory available in whichever domain ran out. That domain is the whole machine for a system-wide OOM and the cgroup for a cgroup OOM, which matters more than it sounds; see attempt 4.
- Plus an adjustment,
oom_score_adj, in/proc/<pid>/oom_score_adj, ranging from-1000to+1000, scaled into the same units before it’s added.
So in plain terms: bigger processes get killed first, but you can bias the
decision either way. -1000 is a special case rather than merely a big
discount: the kernel checks for it up front and drops the task out of victim
selection entirely, which is how things like sshd or critical daemons opt
out.
The kill itself is a SIGKILL. There is no graceful shutdown, no chance to
flush, no atexit handlers. The process is gone.
Attempt 4: scope the whole thing to a cgroup
When you run inside a memory-limited cgroup, there’s a second OOM situation: the cgroup hits its limit even though the host has plenty of RAM. That’s a cgroup OOM, and the kernel runs the same kind of selection — but only over processes inside that cgroup. Your container gets killed; the host doesn’t notice.
Whether “your container” means one process or all of them is a cgroup v1
vs v2 difference worth carrying. Under v1, the kernel picks the worst-scored
process in the cgroup and kills just that one — the container survives, minus
a process. cgroup v2 added memory.oom.group, which tells the kernel to treat
the cgroup as one indivisible workload: every task in it and its descendants
is killed together, or none is. Kubernetes sets that flag by default on
cgroup v2 nodes as of v1.32 (the kubelet’s singleProcessOOMKill option turns
it back off), on the reasoning that a container half-killed is worse than a
container cleanly dead — a half-killed one keeps its pod “Running” while
quietly not working.
This is why a Kubernetes pod with resources.limits.memory: 4Gi can OOM at
4 GB on a node with 256 GB free. The host has memory; the cgroup doesn’t.
It’s also the reason a 137 and a calm node graph are not a contradiction:
the shortage that killed you was measured against a limit somebody typed
into a YAML file, not against the machine, so the node’s memory graph can
look perfectly flat at the moment your container died.
Show the seams
A few things the textbook version skips:
- Overcommit is configurable.
vm.overcommit_memoryhas three modes: heuristic (the default), always (lie aggressively), and never (attempt 1 above — refuse allocations that wouldn’t fit). Knowing the knob exists is useful; flipping it on a general-purpose box usually isn’t. - The killer’s choice can be surprising. The biggest process is often the
one doing the work you care about — your database, your model server. The
kernel doesn’t know that. Tools like
systemd-oomdand policy viaoom_score_adjexist precisely because the default heuristic is naive about importance. - OOM != “no free memory.” Linux deliberately keeps memory full —
unused RAM is wasted RAM, so the page cache will eat anything free. “Free”
memory in
topbeing near zero is normal. OOM only fires when allocation truly fails after reclaim. - It can also kill the wrong cgroup-mate. Where per-process selection is in force — cgroup v1, or v2 with group kill disabled — the killer picks the worst-scored process in that cgroup, which isn’t always the leaker; sometimes it’s the innocent neighbor that just happens to be larger. Group kill trades that failure for a blunter one: nobody innocent is singled out because everybody goes.
dmesgis where the evidence lives. When the OOM killer fires, it logs a whole dump: which process, the score, the memory state at the time. If you’ve ever debugged a mysterious “my container just disappeared,” that log is the breadcrumb.
You started with OOM killer = "system is out of memory" trigger + heuristic that scores processes + SIGKILL on the highest scorer. What did
exit code 137 add to that line? — + the trigger fires when a page is really needed, not at malloc. Usually that moment is your first touch of the
memory. That’s the whole reason it’s a killer and not an error code: by the
time the shortage is real, the allocation that caused it succeeded long ago
and there is no caller left to tell. The scoring heuristic is the kernel
admitting it now has to guess who mattered least, with only size to go on.
Linux’s own strict mode (vm.overcommit_memory=2) shows the alternative —
refuse the allocation up front and let malloc fail honestly — and almost
nobody runs it, which is the most direct evidence of what the trade is worth.
Linux’s bet is that overcommit plus a reaper gets
better aggregate throughput, even though the tail experience — your job
disappearing with a 137 — is uglier.
Check yourself
Before you go — your model server is killed with 137 on a node whose memory graph shows 200 GB free the entire time. What’s the most likely explanation, and where would you look first?
Answer
A cgroup OOM, not a host OOM. The container hit its own limits.memory
cap, and the kernel ran selection over the processes inside that cgroup
only — so the host never came under pressure and its graph stayed flat.
Look at the pod’s memory limit versus its actual working set, and at
dmesg (or the node’s kernel log) for the OOM dump, which names the cgroup
and the chosen victim. Check whether the whole container died or just one
process inside it — on a cgroup v2 node with group kill on, it’s all of them,
which is a different debugging story than one process vanishing. The general
lesson: “out of memory” is always relative to some memory domain, and the
interesting question is which domain.
And a trade-off: your critical database gets OOM-killed on a shared box, so
you set its oom_score_adj to -1000. What have you actually bought, and
what have you moved?
Answer
You’ve dropped that process out of victim selection — but you haven’t created any memory. The shortage is unchanged; the kernel will now pick the next-highest scorer, which might be the thing that fills the database’s queue, or something whose death silently degrades the database anyway. Immunity is a ranking knob, not a capacity knob. If the box is genuinely undersized, all you’ve chosen is who finds out.
Famous related terms
- Overcommit —
overcommit = "promise more memory than you have" + "hope nobody calls the bet"— the policy that makes the OOM killer necessary. - cgroup memory limit —
cgroup memory limit ≈ per-process-group RAM cap— creates a second, narrower OOM domain inside a host. oom_score_adj—oom_score_adj = per-process bias on who gets killed first— the knob for “please don’t kill my database.”systemd-oomd—systemd-oomd ≈ userspace OOM killer that acts before the kernel's does— uses pressure signals to kill earlier and more selectively.- PSI (Pressure Stall Information) —
PSI ≈ "how stalled is the system on memory/CPU/IO right now"— the modern signalsystemd-oomdreads to act before the kernel has to.
Going deeper
- The kernel source,
mm/oom_kill.c— the primary source, and the only way to answer “what exactly does the scoring do today” without trusting a blog post about a formula that has changed more than once. man 5 proc— the reference for whatoom_score,oom_score_adj, and thevm.overcommit_*knobs actually mean before you start tuning them.- The kernel’s cgroup v2 admin guide,
memory-controller section — the primary source for the second OOM domain:
memory.maxvsmemory.high,memory.oom.group, and thememory.eventscounters that tell you an OOM happened at all. - Chris Down’s writing on
systemd-oomdand PSI — the rabbit hole: why you might want to kill something in userspace before the kernel is cornered, and what signal tells you it’s time.