Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why does continuous batching exist?

Static batching works fine for image classifiers and breaks immediately for LLMs. The problem isn't the batch — it's that generation lengths vary, and the slowest sequence holds the GPU hostage.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never run a GPU serving stack. The prose below fills in the seams the pictures skip.

1 · The problem

Row 3, one backpack, twelve minutes in the aisle.

the cabin 3 the doors don’t open until the whole cabin is sorted out same shape one GPU batch, eight users “is this valid JSON? yes/no” 5 tokens “write me a 2,000-token essay” six requests somewhere in between nobody leaves until the essay is finished Your wait is set by the slowest passenger, not by your own request. batching itself is a good idea — it is why serving is affordable at all. the trouble is what happens when the trips take different lengths.
Batching divides the fixed cost of dragging the model’s weights through memory across N requests. For an image classifier, where everyone produces one answer, that is a clean deal. For generation it falls apart, because different requests emit different numbers of tokens.

2 · What static batching wastes

Seven slots finished. The GPU keeps grinding them anyway.

decode iterations → yes/no slot 2 slot 3 slot 4 slot 5 essay solid = real work dashed = slot finished, still occupied, still multiplying weights against padding everyone returns here And new requests queue behind all of it, including the empty slots.
The batch has a fixed tensor shape, so a finished sequence cannot simply leave — its slot keeps producing something until the whole batch is done. On realistic length distributions this is one of the largest sources of wasted GPU in a generation pipeline. Bar lengths are illustrative.

3 · The fix

Schedule per iteration, not per batch.

THE BATCH re-examined after every single iteration free slot emitted EOS → evicted response streams back to its user immediately a waiting request is promoted into the gap, next iteration then run the next iteration on the new composition, and look again out in
The batch becomes a revolving door rather than a fixed cohort. Nobody waits for the car to fill or empty. Note the null case: if every request happens to want the same number of tokens and they all arrive together, this reduces to static batching and buys nothing — the win comes from variance in generation length.

4 · Why it isn’t just “smaller batches”

The two phases want opposite things from the hardware.

prefill — a new arrival thousands of prompt tokens, all at once compute-bound tensor cores near peak — each weight reused across every prompt token decode — everyone already running one token at a time, against a growing KV cache memory-bandwidth-bound compute units half idle while weights stream out of memory put them in the same iteration and the prefill dominates its wall-clock one long prefill … thirty ongoing decodes, all waiting out that one bar So a freed slot poses a question, not just an opportunity. chunked prefill splits the long bar into pieces; disaggregated serving moves the two phases onto separate GPU pools entirely
This is the prefill stalls decode problem, and it is why continuous batching is a family of schedulers rather than a single algorithm. Evicting finished sequences is the easy half that anyone would invent; deciding what to admit is the research area.

5 · The bill for reshuffling

A slot is not free. It owns memory that keeps growing.

GPU memory, after the weights are loaded req A req B req C req D dashed gaps: too small for the next request, too many to ignore every running sequence owns a KV cache that grows with every token it emits so admitting one more request is a promise to store its keys and values for however long it decides to talk Memory, not slot count, is what caps the revolving door.
Once you reshuffle sequences in and out of slots, fragmentation eats real memory and the batch size you can sustain shrinks with it. That is what PagedAttention attacks — reducing fragmentation is what turns iteration-level scheduling into a throughput win rather than a memory crisis.

6 · Keep this card

The whole thing on one index card.

continuous batching = schedule every iteration, not every batch + reshuffle who is in it as sequences finish + a prefill / decode mixing policy ∴ the third line is why this is a research area, not a patch
Picture to keep: a revolving door instead of an elevator. Nobody waits for the car to fill or empty — each person steps out the moment they reach their floor, and the next person in the lobby steps into the gap on the very next rotation. Where it breaks: a revolving door moves one person at a time, while the GPU still processes the whole batch in parallel each rotation. What changes is who is in it.

Why it exists

You’ve had the airport version of this: the plane lands, you’re in row 3 with nothing but a backpack, and you still stand in the aisle for twelve minutes because the doors don’t open until the whole cabin is sorted out. Your trip took as long as the slowest passenger’s.

That is exactly what happens to a one-word question when it shares a GPU batch with an essay request — the running example for this post. Picture eight users hitting the same model. One asks a yes/no question. One asks for a 2,000-token essay. Six want something in between.

Batching itself is a good idea, and it’s why serving is affordable at all: stack N requests into one tensor, hand it to the GPU, and the fixed cost of dragging the model’s weights through memory gets divided by N. For an image classifier that emits one answer per image, that’s a clean deal — everyone arrives and leaves together.

For an LLM it falls apart, for a reason mundane enough to be embarrassing: different requests generate different numbers of tokens. With static batching — pick the eight, run them together, return when they’re all done — the yes/no user waits for the essay-writer. The GPU keeps decoding the batch iteration after iteration, even after seven of the eight sequences have emitted their EOS token and have nothing left to do. Their slots are still “computing” — multiplying weights against padding — because that’s how fixed-shape tensor execution works.

So you’ve pinned the latency of every fast request to the slowest one in the batch. New requests arriving meanwhile queue behind the whole batch, even though half its slots are doing fake work. On workloads with realistic length distributions this isn’t a rounding error; it’s one of the largest sources of wasted GPU in a generation pipeline.

Why it matters now

Every major LLM inference stack does some form of this. vLLM, TensorRT-LLM, TGI, SGLang, llama.cpp’s server mode — the schedulers differ in detail, but the core idea is the same: don’t wait for a batch to finish; reshuffle the batch every iteration.

It matters now because LLM serving economics live and die on aggregate throughput. Hosted providers price per million tokens. Their margin is the gap between what they bill you and what their GPUs cost per token produced. Continuous batching is one of the handful of optimizations that move that gap meaningfully. The Anyscale write-up reports up to 23× throughput with vLLM over naive/static batching while also reducing p50 latency; TGI’s continuous batching lands well short of that in the same comparison (that comparison is read off their chart rather than stated as a headline figure). The original Orca paper reports up to 36.9× throughput at the same latency vs NVIDIA FasterTransformer on a GPT-3 175B workload. The multiplier you see in practice depends heavily on workload, baseline, model, and hardware — long, varied generation lengths give the biggest wins; short uniform ones give you almost nothing — so treat any single number as a data point, not a guarantee.

It also matters because it shapes the public APIs you use. Streaming responses, request-level cancellation, and mixing long and short prompts in the same deployment all benefit substantially from per-iteration scheduling. They don’t strictly require it, but without it the behavior of “how does this request feel?” depends on whoever else happened to be in your batch.

The short answer

continuous batching = iteration-level scheduling + per-token batch reshuffling

Picture to keep: a revolving door instead of an elevator. Nobody waits for the car to fill or empty — each person steps out the moment they arrive at their floor, and the next person in the lobby steps into the gap on the very next rotation. The analogy breaks in one place worth naming: a revolving door moves one person at a time, whereas the GPU still processes the whole batch in parallel each rotation. What changes is who is in it, not that they take turns.

Instead of treating a batch as a fixed group of requests that runs to completion, treat each forward pass through the model as the unit of scheduling. After every iteration, look at what’s in the batch: any sequence that finished gets evicted, freeing its slot; any waiting request can be slotted in for the next iteration. The batch is a revolving door, not a fixed cohort. The GPU spends its cycles on sequences that still have tokens to produce, instead of grinding through padding for ones that already emitted their EOS.

How it works

Static batching is structured like this:

  1. Wait for a batch of N requests.
  2. Run the prompts through the model (prefill).
  3. Decode tokens autoregressively, one iteration per token, for the whole batch, until every sequence has hit EOS or the max-length limit.
  4. Return all N responses.

The pain is step 3. Our yes/no user needs maybe 5 tokens; the essay-writer needs 2,000. After 5 iterations the yes/no sequence is done — but the loop keeps running for 1,995 more, and its slot has to keep producing something (usually a no-op masked-out forward pass) because the tensor has fixed shape. That user waits out the entire essay, and any new request arriving during those 1,995 iterations waits until the whole batch finishes before it can even start. Row 3, backpack, twelve minutes in the aisle.

Continuous batching restructures step 3:

  1. After each iteration of the decode loop, the scheduler looks at the running batch.
  2. Any sequence that just emitted EOS is removed; its response gets streamed back to its user immediately, and its slot becomes free.
  3. Any waiting request can be promoted into a free slot. (In some schedulers, a fresh request gets its prefill done first, possibly interleaved with decode steps from the existing batch — this is where the implementations diverge.)
  4. Run the next iteration on the new batch composition. Loop.

This is iteration-level scheduling. The original term comes from the 2022 Orca paper at OSDI (Yu, Jeong, Kim, Kim, Chun — Orca: A Distributed Serving System for Transformer-Based Generative Models), which introduced both this idea and a complementary one called selective batching — the observation that some operators in a transformer layer (like attention, where each request has its own KV cache) didn’t batch cleanly across sequences of different lengths in their setting, so you batch the parts you can (the matrix multiplies for the linear projections) and run the parts you can’t (attention) per-sequence. The name that ended up in common use is “continuous batching” rather than Orca’s own terminology; the coinage has no clearly documented origin.

Why this isn’t just “smaller batches”

The first time you hear the pitch — “evict finished sequences, slot in new ones” — you might think it’s just dynamic batch sizing. It is more than that. The non-obvious part is that prefill and decode have wildly different costs and arithmetic profiles, and continuous batching has to make a call every iteration about how to mix them.

A new request arriving has a long prompt to process: maybe hundreds or thousands of tokens, all in parallel. That’s a prefill step, and it’s heavy and compute-bound — the GPU finally gets to use its tensor cores near peak because each weight is reused across all the prompt tokens.

A request that’s been around for a while is in the decode phase: one new token at a time, against a growing KV cache. Decode is memory-bandwidth-bound; the compute units sit half-idle while weights stream from HBM.

If you naively let prefill and decode share an iteration, the prefill of a long prompt will dominate the wall-clock of that iteration, and every already-running sequence sees a latency spike on the token they were mid-decoding. This is the classic prefill stalls decode problem. Modern schedulers handle it in a few ways:

So “continuous batching” in 2026 isn’t a single algorithm; it’s a family of schedulers built around iteration-level granularity, all wrestling with the same prefill/decode tension. The Orca paper laid the groundwork; vLLM’s PagedAttention paper (Kwon et al., SOSP 2023) (ACM proceedings) pushed the practical ceiling much higher by attacking KV-cache fragmentation — once you start reshuffling sequences in and out of slots, fragmentation eats real memory, and the max batch size you can sustain shrinks with it.

What you actually get

The headline number people quote is “throughput at iso-latency” — how many tokens per second the system can deliver across all users while keeping each user’s per-token latency under some target. Continuous batching wins on this axis for two reasons:

The wins are biggest when generation lengths are variable and arrival rates are bursty — i.e. when static batching’s worst-case behavior gets triggered constantly. On a workload where every request happens to want exactly the same number of tokens and arrives in lockstep, continuous batching reduces to static batching, and the gap closes.

Where the seams show

A few honest caveats:

You started with continuous batching = iteration-level scheduling + per-token batch reshuffling. What did this post add that makes it hard rather than obvious? — + a prefill/decode mixing policy. Evicting finished sequences is the easy half and anyone would invent it. The reason continuous batching is a research area rather than a patch is that the moment a slot frees up, the scheduler has to decide whether to fill it with a compute-heavy prefill that will stall everyone else’s next token.

Check yourself

Before you go — a colleague benchmarks your serving stack and reports that continuous batching gave them a 1.1× throughput win, nothing like the numbers in the papers. Their workload: a document-classification service where every request is a long document and every response is a single label token. Is the benchmark wrong?

Answer

Probably not — that workload is close to continuous batching’s null case. Every sequence finishes after roughly the same number of decode steps (one), so there’s almost no padding waste and almost no head-of-line blocking to recover: static batching was already near optimal on the decode side. What their workload is dominated by is prefill, which is compute-bound, so the interesting question for them isn’t batch reshuffling at all — it’s prefill throughput and chunking. The rough rule: continuous batching’s win depends far more on the variance in generation length than on raw traffic volume.

And one more: the post says a freed slot can be filled on the next iteration. Why can’t a serving engine just always keep every slot full?

Answer

Because a “slot” isn’t free — each running sequence owns a growing KV cache in GPU memory, and that memory is the real constraint, not the slot count. Admitting one more request means committing to store its keys and values for however long it generates, and if you over-admit you run out of memory mid-generation and have to preempt or recompute someone. That’s precisely why PagedAttention mattered: reducing fragmentation raises how many sequences you can safely keep resident, which is what turns iteration-level scheduling into a throughput win rather than a memory crisis.

Going deeper