Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why garbage collectors pause your program

A tracing collector can't safely move or free an object while your code is mid-read. Freezing every application thread is the obviously correct answer — and generations, write barriers and concurrent marking are all ways of making that freeze shorter without giving up what it buys.

Systems intermediate May 14, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Six pictures for a reader who has never thought about where memory goes, following one thing you have felt: the stutter in a scrolling feed.

1 · The problem

The program stopped, and then carried on.

frames going out, one every 16 milliseconds nothing the hitch you felt Nothing crashed, and the program did not do more work. something else was using that time: the part of the runtime that works out which objects are dead and takes their memory back It froze your threads on purpose, and the reason is correctness. a pause barely moves your average — it wrecks your tail, and the tail is the number your users actually feel
The stutter is not the runtime being sloppy. It is a deliberate freeze, bought on purpose — and the rest of these pictures are about what it buys and how short it can be made.

2 · Why a freeze at all

You cannot take inventory while the forklifts are moving.

the collector walks from the roots, follows every pointer, marks what it reaches as alive your code at the same instant, rewriting those same pointers as it runs so the collector sees a graph that never existed — half old, half new It frees an object you are about to use, or misses one you just saved. use-after-free and a corrupted heap: precisely the bugs a managed language promised to abolish the obviously correct fix: stop every application thread, then look a consistent view of the heap and a running program are in tension, and correctness wins
This is the whole reason the pause exists. The collector and your code are reading and rewriting the same graph, and moving an object is worse still — every pointer to it has to be corrected before anyone follows the old address.

3 · The bet that shrinks it

Most objects die young, so collect the young part often.

young generation loop temporaries, request buffers, intermediate strings — nearly all dead survivors get promoted old generation caches and long-lived state collected rarely One giant pause becomes many short ones over a small region. trace and copy the handful of survivors, then declare the whole rest of that space free in one move that is the difference between a visible stutter once a minute and a dropped frame you never notice not every collector takes this bet — Go’s is non-generational, and ZGC only added generations years in
Generational collection is an empirical bet, not a law: it pays off exactly to the extent that your allocations really do die young. It makes the freeze small and frequent rather than rare and enormous.

4 · Where the bill goes

To collect the young part alone, your own pointer stores pay a tax.

an old object a brand-new object points at that pointer is a root the young-gen collection would otherwise never find and scanning the whole old generation to look for it would erase the savings So the runtime notes the store at the moment it happens. a few extra instructions run on your pointer writes, recording old-pointing-at-young for the collector to read later the pause did not vanish — part of it moved into your own code the same family of trick makes concurrent marking safe, so the marker cannot miss a pointer you rewrite behind it
Follow the money and the shape of the whole subject appears. The cheap young-gen collection is paid for by a standing tax on every pointer store in your program — cost moved, not removed.

5 · How small it gets

Sub-millisecond, and no longer growing with your heap.

mark the graph while the program runs, stopping the world only briefly to reconcile move objects while it runs too, with a barrier steering each access to the right copy and run the phases that do stop the world on many collector threads at once ZGC’s stated goal is pauses under a millisecond that don’t grow with the heap. because the phases that still stop the world no longer scan the heap — a different regime, not a smaller version of the old one what it is not is zero: concurrent work competes with your program for cores and bandwidth, and read barriers tax your loads you bought predictability with throughput, which is the right trade for a latency-bound service and the wrong one for a batch job
The freeze got small and stopped scaling — which is not the same as going away. Every “low-pause” collector is a decision about where to put the bill, and someone always pays it.

6 · Keep this card

The whole thing on one index card.

stop-the-world pause = freeze every application thread… … so the collector sees a heap that isn’t changing ∴ the cost never disappears — it only moves generations move it into rarer, smaller freezes barriers move it into your pointer stores; concurrency moves it onto spare cores so the question was never “can we avoid pausing” — it is how short, how predictable, and how much throughput you will trade
Picture to keep: taking inventory of a warehouse while forklifts are still moving pallets around. You halt the forklifts, count, and wave them back in. Where it breaks: the collector isn’t only counting — it hauls pallets itself when it compacts, which is why “just let the forklifts keep working” is so much harder than it sounds.

Why it exists

You’re playing a game, or scrolling a feed, and every few seconds it hitches — a tiny stutter, a dropped frame, a moment where the input lag spikes and then clears. Nothing crashed. The program didn’t do more work. It just… stopped, briefly, and then carried on. If the runtime is garbage-collected — Java, Go, C#, JavaScript, Python — one possible culprit is the garbage collector doing its job.

Here’s the problem it’s solving. Your program allocates objects constantly and almost never explicitly frees them — that’s the whole point of a managed language. So something has to figure out which objects are dead (nothing points to them anymore) and reclaim their memory. The dominant approach is a tracing collector: it walks the graph of live objects, starting from the roots, follows every pointer, marks everything it reaches as alive, and treats the rest as garbage. (The other major family, reference counting, works differently — more on that at the end. This post is about tracing collectors, which is where the classic stop-the-world pause comes from. Refcounted runtimes aren’t pause-free either; their stalls just have different shapes, like a cascade of frees when one big object graph loses its last reference.)

But the program is also running. Your code is following those same pointers, reading fields, writing new pointers into objects. If the collector is reading the object graph at the same moment your code is rewriting it, the collector sees a graph that never actually existed — half-old, half-new — and can free an object you’re about to use, or miss an object you just made reachable. That’s a use-after-free or a corrupted heap, the exact bugs managed languages promised to abolish.

The simplest fix that is obviously correct: stop the program. Freeze every application thread (the runtime calls these threads the mutators), take your snapshot, do the work, resume. (Real collectors need something subtler than a literal frozen photograph of the heap — the concurrent ones below get by with carefully maintained invariants instead — but the freeze is the version that’s obviously correct, which is why it came first.) That freeze is the stop-the-world pause. It exists because a consistent view of the heap and a running program are fundamentally in tension — and correctness wins.

Why it matters now

A pause is easy to live with when “the program” is a batch job — total throughput is what you’re measuring, and a freeze every so often barely moves it. It’s much harder to live with when the program is a request server with a p99 latency budget, a game holding a 16-millisecond frame deadline, or a trading system where a 50-millisecond hiccup is a real loss. A GC pause barely moves your average — it wrecks your tail. One unlucky request in a thousand eats the full pause, and that’s the number your users and your SLOs actually feel.

That tension is why so much modern runtime engineering is, specifically, pause engineering. Go made low pause time an explicit design goal. The JVM ships multiple collectors — G1, and the newer low-latency ZGC and Shenandoah — that exist largely to shrink or break up the stop-the-world window. None of them eliminate it entirely. The interesting question was never “can we avoid pausing” — it’s “how short, how predictable, and how much throughput do we trade to get there.”

The short answer

stop-the-world pause = "freeze all mutator threads" + "so the collector sees a heap that isn't changing under it"

Picture to keep: taking inventory of a warehouse while forklifts are still moving pallets around. You can’t count what’s moving, so you halt the forklifts, count, and wave them back in. Where the picture breaks: the collector isn’t only counting — it’s also hauling pallets itself when it compacts, which is why “just let the forklifts keep working” is so much harder than it sounds.

You can’t safely move or free an object while another thread might be reading it. The bluntest way to guarantee that is to make sure no thread is reading anything. The collector freezes the mutators, gets a stable snapshot, does the dangerous part, and resumes them. The clever machinery in modern tracing collectors — generations, concurrent marking, read/write barriers — is an effort to make that frozen window smaller, rarer, or more predictable, without giving up the correctness the freeze buys.

How it works

Start with the naive collector and watch where the pause comes from.

Mark. From the roots, walk every reachable pointer and mark each object live. Sweep (or compact). Reclaim everything unmarked; optionally move the survivors so they sit contiguously and defeat fragmentation. With no barriers and nothing stopped, the mark phase is wrong the instant a mutator rewrites a pointer mid-walk. The compaction phase is even more delicate: if the collector moves an object, every pointer to it must be updated, and a mutator must not dereference the old address in the gap between “object copied” and “pointers fixed.” Both phases want the world stopped.

So why isn’t every pause enormous? Because of one empirical observation about how many programs behave:

The generational hypothesis: most objects die young. In many workloads, the bulk of allocations — loop temporaries, request-scoped buffers, intermediate strings — become garbage very quickly. A smaller set (caches, long-lived state) survives a long time.

Generational GC turns that observation into a strategy. Split the heap into a young generation and an old generation. New objects are born in the young gen. Most of them die there. So collect the young gen often — and in the common copying design a young-gen collection is cheap, because you only trace and copy the handful of survivors, then declare the entire rest of that space free in one move. Objects that survive a few young-gen collections get promoted (tenured) into the old gen, which you collect rarely.

Not every collector takes this bet. Go’s collector, notably, is non-generational and non-compacting — it relies on concurrent marking instead (more below). And the traffic runs both ways: ZGC shipped without generations, added a generational mode in JDK 21, made it the default in JDK 23, and removed the non-generational one in JDK 24. Generational GC is a common pause-reduction strategy, not a universal one.

The payoff: instead of one giant pause to trace the whole heap, you get frequent short pauses over a small region, and the expensive full-heap collection happens rarely. Same total correctness, with the pause cost spread into small predictable chunks. That’s the difference between your feed stuttering visibly once a minute and dropping a frame you never notice — the freezes didn’t stop, they got small enough to hide inside the gaps.

There’s a catch, and it’s where the seam shows. To collect the young gen alone, you need its roots — and an old-gen object holding a pointer into the young gen is a root you’d otherwise miss. The runtime can’t afford to scan the whole old gen to find those pointers; that would erase the savings. So it tracks them as they’re created, using a write barrier: a small piece of code injected on pointer-writes in your program that records when an old-gen object starts pointing into the young gen. The collector’s record of those cross-region references is called a remembered set; a common way to maintain one cheaply is a card table, which just marks the coarse region (“card”) a write landed in, so the collector rescans a small strip of the old gen rather than all of it. (Names and implementations vary by runtime; the idea — note the cross-generation write when it happens, not later — is the durable part.) You pay a few instructions on pointer stores so that young-gen collections stay cheap. The pause didn’t vanish — part of its cost was amortized into the mutator.

Shrinking the pause further

Generational GC makes pauses small but doesn’t remove them. The next moves:

Worth putting a number on how far that has been pushed: ZGC’s stated design goal is pauses typically under a millisecond, and — the part that matters more — pauses that don’t grow as the heap does, because the phases that still stop the world don’t scan the heap. That is a different regime from a collector whose pause target is a knob you set in the hundreds of milliseconds. What it isn’t is zero: the freeze got small and stopped scaling with your heap, which is not the same as going away.

Every one of these trades something. Concurrent work competes with your program for CPU and memory bandwidth, so you lose throughput to gain predictability. Write barriers add a small cost to pointer writes; the load/read barriers in concurrent-compacting collectors add cost to pointer reads. There is no free collector — only different points on the pause-versus-throughput curve, and the right pick depends on whether you’re running a batch job or a latency-bound service.

So the stutter in your feed was never the collector doing something wrong. You started with stop-the-world pause = freeze all mutator threads + so the collector sees a heap that isn't changing under it. What did the mechanism add? — + the cost never disappears, it only moves. Generations move it into rarer, smaller freezes; write barriers move part of it into your own pointer stores; concurrent marking moves it onto spare cores. Every “low-pause” collector is a decision about where to put the bill, and someone always pays it.

Check yourself

Before you go — a service moves to a low-pause collector and its p99 latency improves a lot. Throughput, measured in requests per second at saturation, gets slightly worse. Did something go wrong?

Answer

No — that’s the trade working as designed. The pause didn’t vanish; most of the collection work moved to running alongside your threads. That work now competes for the same cores and memory bandwidth your request handlers want, and the barriers on pointer operations cost a few instructions each, all the time, in your code. You bought predictability with throughput. The question to ask isn’t “which collector is faster” but “which resource do I have spare” — a batch job with a throughput target should probably make the opposite choice.

And one more — a cache holds a million long-lived objects. Would you expect that to make young-generation collections slower, faster, or neither?

Answer

Neither, mostly — if the cache is quiet. Young-gen collections trace only the young generation plus its roots, so a million untouched old-gen objects are simply not visited; they’re close to free, not literally free (they still occupy memory, which affects how often the old gen has to be collected). But start writing into that cache — storing freshly allocated request objects into long-lived entries — and each such store is an old-to-young pointer the write barrier has to record, growing the remembered set the next young-gen collection must scan as roots. So the size of the cache is nearly free; the rate at which old objects are made to point at new ones is what shows up in your pause times. That’s a decent example of running the model: the answer comes from the mechanism, not from a rule of thumb about cache sizes.

Going deeper