Why does continuous batching exist?
Static batching works fine for image classifiers and breaks immediately for LLMs. The problem isn't the batch — it's that generation lengths vary, and the slowest sequence holds the GPU hostage.
On this page
The picture version
Six pictures for a reader who has never run a GPU serving stack. The prose below fills in the seams the pictures skip.
1 · The problem
Row 3, one backpack, twelve minutes in the aisle.
2 · What static batching wastes
Seven slots finished. The GPU keeps grinding them anyway.
3 · The fix
Schedule per iteration, not per batch.
4 · Why it isn’t just “smaller batches”
The two phases want opposite things from the hardware.
5 · The bill for reshuffling
A slot is not free. It owns memory that keeps growing.
6 · Keep this card
The whole thing on one index card.
Why it exists
You’ve had the airport version of this: the plane lands, you’re in row 3 with nothing but a backpack, and you still stand in the aisle for twelve minutes because the doors don’t open until the whole cabin is sorted out. Your trip took as long as the slowest passenger’s.
That is exactly what happens to a one-word question when it shares a GPU batch with an essay request — the running example for this post. Picture eight users hitting the same model. One asks a yes/no question. One asks for a 2,000-token essay. Six want something in between.
Batching itself is a good idea, and it’s why serving is affordable at all: stack N requests into one tensor, hand it to the GPU, and the fixed cost of dragging the model’s weights through memory gets divided by N. For an image classifier that emits one answer per image, that’s a clean deal — everyone arrives and leaves together.
For an LLM it falls apart, for a reason mundane enough to be embarrassing: different requests generate different numbers of tokens. With static batching — pick the eight, run them together, return when they’re all done — the yes/no user waits for the essay-writer. The GPU keeps decoding the batch iteration after iteration, even after seven of the eight sequences have emitted their EOS token and have nothing left to do. Their slots are still “computing” — multiplying weights against padding — because that’s how fixed-shape tensor execution works.
So you’ve pinned the latency of every fast request to the slowest one in the batch. New requests arriving meanwhile queue behind the whole batch, even though half its slots are doing fake work. On workloads with realistic length distributions this isn’t a rounding error; it’s one of the largest sources of wasted GPU in a generation pipeline.
Why it matters now
Every major LLM inference stack does some form of this. vLLM, TensorRT-LLM, TGI, SGLang, llama.cpp’s server mode — the schedulers differ in detail, but the core idea is the same: don’t wait for a batch to finish; reshuffle the batch every iteration.
It matters now because LLM serving economics live and die on aggregate throughput. Hosted providers price per million tokens. Their margin is the gap between what they bill you and what their GPUs cost per token produced. Continuous batching is one of the handful of optimizations that move that gap meaningfully. The Anyscale write-up reports up to 23× throughput with vLLM over naive/static batching while also reducing p50 latency; TGI’s continuous batching lands well short of that in the same comparison (that comparison is read off their chart rather than stated as a headline figure). The original Orca paper reports up to 36.9× throughput at the same latency vs NVIDIA FasterTransformer on a GPT-3 175B workload. The multiplier you see in practice depends heavily on workload, baseline, model, and hardware — long, varied generation lengths give the biggest wins; short uniform ones give you almost nothing — so treat any single number as a data point, not a guarantee.
It also matters because it shapes the public APIs you use. Streaming responses, request-level cancellation, and mixing long and short prompts in the same deployment all benefit substantially from per-iteration scheduling. They don’t strictly require it, but without it the behavior of “how does this request feel?” depends on whoever else happened to be in your batch.
The short answer
continuous batching = iteration-level scheduling + per-token batch reshuffling
Picture to keep: a revolving door instead of an elevator. Nobody waits for the car to fill or empty — each person steps out the moment they arrive at their floor, and the next person in the lobby steps into the gap on the very next rotation. The analogy breaks in one place worth naming: a revolving door moves one person at a time, whereas the GPU still processes the whole batch in parallel each rotation. What changes is who is in it, not that they take turns.
Instead of treating a batch as a fixed group of requests that runs to completion, treat each forward pass through the model as the unit of scheduling. After every iteration, look at what’s in the batch: any sequence that finished gets evicted, freeing its slot; any waiting request can be slotted in for the next iteration. The batch is a revolving door, not a fixed cohort. The GPU spends its cycles on sequences that still have tokens to produce, instead of grinding through padding for ones that already emitted their EOS.
How it works
Static batching is structured like this:
- Wait for a batch of N requests.
- Run the prompts through the model (prefill).
- Decode tokens autoregressively, one iteration per token, for the whole batch, until every sequence has hit EOS or the max-length limit.
- Return all N responses.
The pain is step 3. Our yes/no user needs maybe 5 tokens; the essay-writer needs 2,000. After 5 iterations the yes/no sequence is done — but the loop keeps running for 1,995 more, and its slot has to keep producing something (usually a no-op masked-out forward pass) because the tensor has fixed shape. That user waits out the entire essay, and any new request arriving during those 1,995 iterations waits until the whole batch finishes before it can even start. Row 3, backpack, twelve minutes in the aisle.
Continuous batching restructures step 3:
- After each iteration of the decode loop, the scheduler looks at the running batch.
- Any sequence that just emitted EOS is removed; its response gets streamed back to its user immediately, and its slot becomes free.
- Any waiting request can be promoted into a free slot. (In some schedulers, a fresh request gets its prefill done first, possibly interleaved with decode steps from the existing batch — this is where the implementations diverge.)
- Run the next iteration on the new batch composition. Loop.
This is iteration-level scheduling. The original term comes from the 2022 Orca paper at OSDI (Yu, Jeong, Kim, Kim, Chun — Orca: A Distributed Serving System for Transformer-Based Generative Models), which introduced both this idea and a complementary one called selective batching — the observation that some operators in a transformer layer (like attention, where each request has its own KV cache) didn’t batch cleanly across sequences of different lengths in their setting, so you batch the parts you can (the matrix multiplies for the linear projections) and run the parts you can’t (attention) per-sequence. The name that ended up in common use is “continuous batching” rather than Orca’s own terminology; the coinage has no clearly documented origin.
Why this isn’t just “smaller batches”
The first time you hear the pitch — “evict finished sequences, slot in new ones” — you might think it’s just dynamic batch sizing. It is more than that. The non-obvious part is that prefill and decode have wildly different costs and arithmetic profiles, and continuous batching has to make a call every iteration about how to mix them.
A new request arriving has a long prompt to process: maybe hundreds or thousands of tokens, all in parallel. That’s a prefill step, and it’s heavy and compute-bound — the GPU finally gets to use its tensor cores near peak because each weight is reused across all the prompt tokens.
A request that’s been around for a while is in the decode phase: one new token at a time, against a growing KV cache. Decode is memory-bandwidth-bound; the compute units sit half-idle while weights stream from HBM.
If you naively let prefill and decode share an iteration, the prefill of a long prompt will dominate the wall-clock of that iteration, and every already-running sequence sees a latency spike on the token they were mid-decoding. This is the classic prefill stalls decode problem. Modern schedulers handle it in a few ways:
- Chunked prefill — split a long prompt’s prefill into pieces small enough to fit alongside ongoing decodes without dominating the iteration. The technique was named and analyzed in the SARATHI paper (Agrawal et al., 2023); vLLM and others have adopted variants.
- Disaggregated prefill/decode — run the two phases on different GPU pools entirely, paying the cost of shipping the KV cache between them, in exchange for keeping the two phases from stepping on each other’s latency. DistServe (OSDI 2024) is the cleanest reference; it’s a serving pattern in its own right rather than a continuous-batching detail.
- Priority and admission control — sometimes you’d rather hold a new request for one extra iteration than spike the latency of 30 ongoing decodes.
So “continuous batching” in 2026 isn’t a single algorithm; it’s a family of schedulers built around iteration-level granularity, all wrestling with the same prefill/decode tension. The Orca paper laid the groundwork; vLLM’s PagedAttention paper (Kwon et al., SOSP 2023) (ACM proceedings) pushed the practical ceiling much higher by attacking KV-cache fragmentation — once you start reshuffling sequences in and out of slots, fragmentation eats real memory, and the max batch size you can sustain shrinks with it.
What you actually get
The headline number people quote is “throughput at iso-latency” — how many tokens per second the system can deliver across all users while keeping each user’s per-token latency under some target. Continuous batching wins on this axis for two reasons:
- Less padding waste. Slots that would have idled on already- finished sequences get reused for new work instead.
- No head-of-line blocking. A short request doesn’t have to wait for a long one to finish before it gets served — nobody stands in the aisle behind the whole cabin.
The wins are biggest when generation lengths are variable and arrival rates are bursty — i.e. when static batching’s worst-case behavior gets triggered constantly. On a workload where every request happens to want exactly the same number of tokens and arrives in lockstep, continuous batching reduces to static batching, and the gap closes.
Where the seams show
A few honest caveats:
- Per-request latency variance gets weirder, not necessarily smaller. Your tokens-per-second now depends on the iteration-by-iteration composition of the batch, which depends on what every other user is doing. Removing head-of-line blocking should help the tail, but predictability for any single request gets harder to reason about — and there’s no public measurement quantifying the net effect across realistic workloads.
- It interacts with everything. Speculative decoding, prompt caching, KV-cache eviction, MoE expert balancing — all of them now have to be implemented to play with a batch whose membership changes every iteration. A surprising amount of inference-engine engineering is “make feature X coexist with continuous batching.” My read — an inference, not a documented claim — is that this is a big part of why feature parity across engines feels so patchy.
- The headline numbers are workload- and baseline-specific. Anyscale’s 23× was vLLM vs naive/static batching; Orca’s 36.9× was vs FasterTransformer on GPT-3 175B at iso-latency. The right way to read these: “the looser the comparison baseline, the bigger the multiplier reported.” Don’t quote the headline number without a caveat.
You started with continuous batching = iteration-level scheduling + per-token batch reshuffling. What did this post add that makes it hard
rather than obvious? — + a prefill/decode mixing policy. Evicting
finished sequences is the easy half and anyone would invent it. The
reason continuous batching is a research area rather than a patch is
that the moment a slot frees up, the scheduler has to decide whether to
fill it with a compute-heavy prefill that will stall everyone else’s
next token.
Check yourself
Before you go — a colleague benchmarks your serving stack and reports that continuous batching gave them a 1.1× throughput win, nothing like the numbers in the papers. Their workload: a document-classification service where every request is a long document and every response is a single label token. Is the benchmark wrong?
Answer
Probably not — that workload is close to continuous batching’s null case. Every sequence finishes after roughly the same number of decode steps (one), so there’s almost no padding waste and almost no head-of-line blocking to recover: static batching was already near optimal on the decode side. What their workload is dominated by is prefill, which is compute-bound, so the interesting question for them isn’t batch reshuffling at all — it’s prefill throughput and chunking. The rough rule: continuous batching’s win depends far more on the variance in generation length than on raw traffic volume.
And one more: the post says a freed slot can be filled on the next iteration. Why can’t a serving engine just always keep every slot full?
Answer
Because a “slot” isn’t free — each running sequence owns a growing KV cache in GPU memory, and that memory is the real constraint, not the slot count. Admitting one more request means committing to store its keys and values for however long it generates, and if you over-admit you run out of memory mid-generation and have to preempt or recompute someone. That’s precisely why PagedAttention mattered: reducing fragmentation raises how many sequences you can safely keep resident, which is what turns iteration-level scheduling into a throughput win rather than a memory crisis.
Famous related terms
- Iteration-level scheduling —
iteration-level scheduling = treat each forward pass as the scheduling unit + reshuffle the batch each step. The mechanism underneath continuous batching; the name from the Orca paper. - Selective batching —
selective batching = batch the layers you can + run per-sequence the layers you can't— the trick that makes iteration-level scheduling work despite per-sequence attention state. - Static batching —
static batching = pick N requests + run them together until they all finish— the obvious approach; the one continuous batching replaces. - PagedAttention —
PagedAttention = OS-style paging applied to the KV cache— reduces the memory fragmentation that would otherwise cap how much continuous batching can buy you at scale. See vLLM. - Chunked prefill —
chunked prefill = split long prompt prefills into pieces + interleave with ongoing decodes— keeps a single big new request from spiking everyone else’s per-token latency. - Disaggregated prefill/decode —
disaggregated serving = prefill GPUs + decode GPUs + shipped KV cache between them— physical separation of the two phases when even chunked prefill isn’t isolation enough. - KV cache —
KV cache = per-request stored attention K/V tensors + reused on every decode step— the per-request state that makes scheduling messy: every running sequence carries its own growing chunk of GPU memory, and when you evict a sequence you have to reclaim it.
Going deeper
- Yu, Jeong, Kim, Kim, Chun — Orca: A Distributed Serving System for Transformer-Based Generative Models (OSDI 2022) — the primary source, and where to go for “what exactly is iteration-level scheduling, and what does selective batching have to do with it?”
- Anyscale — How continuous batching enables 23x throughput in LLM inference — the explainer, for “what does the padding waste actually look like in a chart, and how much does it cost me?”
- Kwon et al. — Efficient Memory Management for LLM Serving with PagedAttention (SOSP 2023) — the rabbit hole, if the Check yourself memory answer above left you wondering how far you can push slot occupancy before the KV cache breaks you.