Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why Linux has an OOM killer

Linux promises memory it doesn't have, then has to break the promise — the OOM killer is the reaper that decides who dies so the system can keep running.

Systems intermediate Apr 29, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never seen a process disappear without saying anything. The prose below fills in the seams.

1 · The problem

The log just stops. Then a number.

processing batch 41… processing batch 42… processing batch 43… _ and then nothing at all all you get 137 the only clue no stack trace, no exception, no “out of memory” line
Nothing in the program was consulted, so nothing in the program could log it. Whatever killed it did so from outside.

2 · The bet

Linux promises more memory than it has.

malloc(8 GB) kernel: yes malloc(8 GB) kernel: yes malloc(8 GB) kernel: yes each says yes PROMISED 24 GB ACTUALLY ON THE BOARD 8 GB nothing has crashed — nobody has touched most of it yet
Overcommit is a calculated bet: programs reserve far more than they touch, so the kernel says yes to more than exists. Most days the bet pays off.

3 · When the bet loses

The bill arrives long after the sale.

malloc(8 GB) returns a pointer minutes pass buf[0] = 1 the first real touch needs a page KERNEL looks for a page there is none Nobody is left to return an error to. the malloc that lied succeeded minutes ago 1. panic the whole machine everybody dies 2. pick one victim everyone else lives
The shortage is discovered when a real page is needed — usually your first touch — not when it was requested. Linux picks option two, and that reaper is the OOM killer.

4 · Choosing the victim

Score everybody. Kill the worst.

how much memory it occupies, as a share of the domain that ran out log-shipper 40 sidecar 30 model-server 780 sshd oom_score_adj = −1000 → dropped from selection SIGKILL — no cleanup, no last words
The score tracks size, not importance — so the biggest process usually dies, and the biggest process is often the one the machine existed for.

5 · The second kind of “out of memory”

The host is fine. Your container is not.

THE HOST 256 GB, mostly idle host memory graph: flat your container limits.memory: 4Gi full → OOM in here killed 137 cgroup v1 the worst process dies cgroup v2, group kill all of them die together
“Out of memory” is always relative to some memory domain. A cgroup limit is a second, much smaller domain — which is why a 137 and a calm node graph are not a contradiction.

6 · Keep this card

The whole thing on one index card.

OOM KILLER = a promise the kernel can’t keep + the bill arriving at first touch + SIGKILL for whoever scores worst …in whichever domain actually ran out
Picture to keep: an overbooked flight. The airline sold more seats than the plane has. When everyone shows up, the gate agent doesn’t cancel the flight — they bump somebody. The kernel bumps by size, which is often the one passenger the flight existed for.

Why it exists

Your container is gone. No stack trace, no exception, no “out of memory” line in the application log — the last thing it printed was ordinary work, and then nothing. The orchestrator says the process exited with code 137. We’ll follow that container the whole way down.

You probably assume that running out of memory means malloc returns NULL and your program handles it. On Linux, usually not. Your program never found out it was out of memory; something else decided it should stop existing.

Here’s why. A process calls malloc(8 GB) and the kernel says yes. Run a few of these and total “allocated” memory exceeds physical RAM plus swap. Nothing has crashed, because none of those processes have touched most of that memory — malloc reserved address space, not pages.

Then one of them starts writing. The kernel has to find a real, physical page to back each virtual page being touched. At some point it can’t: every page is in use, swap is full, the page cache is already squeezed flat. The kernel made a promise it can’t keep, and — this is the key part — it’s discovering that during a memory access, not during malloc. There’s no error code to return to anyone. The malloc that lied already succeeded, possibly minutes ago.

Two options exist at that moment. One: panic the whole machine — every process dies, the box reboots. Two: pick a victim, kill it, free its pages, and let everyone else live. Linux chose option two and built the OOM killer to do it. Exit code 137 is your container being option two.

The OOM killer exists because Linux deliberately overcommits memory. Overcommit is a calculated bet that programs reserve far more than they touch — and most of the time the bet pays off. The OOM killer is what happens when it doesn’t.

Why it matters now

If you run anything in containers — Kubernetes pods, Docker containers, an ML training job inside a memory-limited cgroup — you have already met the OOM killer. That exit code 137 decodes as 128 + 9: the shell convention for “terminated by signal 9,” and signal 9 is SIGKILL, the one signal a process cannot catch, block, or ignore. Nothing in your program was consulted. Worth knowing the boundary: 137 means something sent SIGKILL — a kill -9, a runtime tearing the container down, an orchestrator timing it out. OOM is the usual suspect in a container, not the only one.

It matters even more now because of large models. An LLM that loads 70 GB of weights, a fine-tuning run that spikes activation memory, a vector index that grows past what you sized for — these are exactly the workloads that will push real pages into existence and force the kernel to make good on its promises. AI infra is OOM-killer infra.

The short answer

OOM killer = "system is out of memory" trigger + heuristic that scores processes + SIGKILL on the highest scorer

Picture to keep: an overbooked flight. The airline sold more seats than the plane has, betting some passengers won’t show. When they all show, the gate agent doesn’t cancel the flight — they pick someone and bump them. Where it breaks: the gate agent bumps by who’s cheapest to bump, whereas the kernel bumps by who’s biggest, which is often the one passenger the flight existed for.

When allocation fails and there’s nowhere left to reclaim from, the kernel walks the process list, computes a “badness” score for each one, and sends SIGKILL to the worst offender. The score roughly tracks how much memory the process is using, with adjustments for privilege and policy.

How it works

Attempt 1: don’t overcommit — fail honestly

The kernel could refuse any malloc it couldn’t back with real RAM plus swap. No lying, no reaper, malloc returns NULL and your program decides what to do.

Why it breaks: two ways. Programs routinely reserve far more than they touch — a fork of a 10 GB process, a sparse array, a thread stack — so honest accounting refuses allocations that would have been fine and leaves real RAM idle. And almost nobody checks malloc’s return value, so “honest” failure mostly manifests as a null-pointer crash at a random location instead of a clean error. Linux still offers this mode (vm.overcommit_memory=2); most systems don’t run it.

Attempt 2: overcommit, and deal with it later

Overcommit takes the bet. Say yes, back pages only when touched, and reclaim aggressively — evict page cache, swap out cold anonymous pages — when things get tight.

Why it breaks: sometimes reclaim isn’t enough. The trigger is a failed page allocation under __alloc_pages that the kernel cannot satisfy even after reclaiming. Now it’s mid-page-fault with nothing to hand back and no caller who can cope. Something has to die.

Attempt 3: kill somebody — but who?

Killing at random works, technically. It also means a one-line log-shipper sidecar can be sacrificed while the 60 GB leaker that caused the crisis survives and triggers the next one thirty seconds later.

So selection uses a score per process, exposed at /proc/<pid>/oom_score. The current formula dates from the badness-heuristic rewrite in Linux 2.6.36, which is also when /proc/<pid>/oom_score_adj replaced the older oom_adj. It is roughly:

So in plain terms: bigger processes get killed first, but you can bias the decision either way. -1000 is a special case rather than merely a big discount: the kernel checks for it up front and drops the task out of victim selection entirely, which is how things like sshd or critical daemons opt out.

The kill itself is a SIGKILL. There is no graceful shutdown, no chance to flush, no atexit handlers. The process is gone.

Attempt 4: scope the whole thing to a cgroup

When you run inside a memory-limited cgroup, there’s a second OOM situation: the cgroup hits its limit even though the host has plenty of RAM. That’s a cgroup OOM, and the kernel runs the same kind of selection — but only over processes inside that cgroup. Your container gets killed; the host doesn’t notice.

Whether “your container” means one process or all of them is a cgroup v1 vs v2 difference worth carrying. Under v1, the kernel picks the worst-scored process in the cgroup and kills just that one — the container survives, minus a process. cgroup v2 added memory.oom.group, which tells the kernel to treat the cgroup as one indivisible workload: every task in it and its descendants is killed together, or none is. Kubernetes sets that flag by default on cgroup v2 nodes as of v1.32 (the kubelet’s singleProcessOOMKill option turns it back off), on the reasoning that a container half-killed is worse than a container cleanly dead — a half-killed one keeps its pod “Running” while quietly not working.

This is why a Kubernetes pod with resources.limits.memory: 4Gi can OOM at 4 GB on a node with 256 GB free. The host has memory; the cgroup doesn’t. It’s also the reason a 137 and a calm node graph are not a contradiction: the shortage that killed you was measured against a limit somebody typed into a YAML file, not against the machine, so the node’s memory graph can look perfectly flat at the moment your container died.

Show the seams

A few things the textbook version skips:

You started with OOM killer = "system is out of memory" trigger + heuristic that scores processes + SIGKILL on the highest scorer. What did exit code 137 add to that line? — + the trigger fires when a page is really needed, not at malloc. Usually that moment is your first touch of the memory. That’s the whole reason it’s a killer and not an error code: by the time the shortage is real, the allocation that caused it succeeded long ago and there is no caller left to tell. The scoring heuristic is the kernel admitting it now has to guess who mattered least, with only size to go on. Linux’s own strict mode (vm.overcommit_memory=2) shows the alternative — refuse the allocation up front and let malloc fail honestly — and almost nobody runs it, which is the most direct evidence of what the trade is worth. Linux’s bet is that overcommit plus a reaper gets better aggregate throughput, even though the tail experience — your job disappearing with a 137 — is uglier.

Check yourself

Before you go — your model server is killed with 137 on a node whose memory graph shows 200 GB free the entire time. What’s the most likely explanation, and where would you look first?

Answer

A cgroup OOM, not a host OOM. The container hit its own limits.memory cap, and the kernel ran selection over the processes inside that cgroup only — so the host never came under pressure and its graph stayed flat. Look at the pod’s memory limit versus its actual working set, and at dmesg (or the node’s kernel log) for the OOM dump, which names the cgroup and the chosen victim. Check whether the whole container died or just one process inside it — on a cgroup v2 node with group kill on, it’s all of them, which is a different debugging story than one process vanishing. The general lesson: “out of memory” is always relative to some memory domain, and the interesting question is which domain.

And a trade-off: your critical database gets OOM-killed on a shared box, so you set its oom_score_adj to -1000. What have you actually bought, and what have you moved?

Answer

You’ve dropped that process out of victim selection — but you haven’t created any memory. The shortage is unchanged; the kernel will now pick the next-highest scorer, which might be the thing that fills the database’s queue, or something whose death silently degrades the database anyway. Immunity is a ranking knob, not a capacity knob. If the box is genuinely undersized, all you’ve chosen is who finds out.

Going deeper