Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do LLM responses stream?

It's not for show. The model literally generates one token at a time, and forcing it to buffer the full answer before sending would make every chat app feel broken. Streaming is the network shape of an autoregressive process.

Networking intro Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Six pictures for a reader who has only ever watched an answer type itself out. The prose below fills in the seams the pictures skip.

1 · The thing you assumed

You probably read the typewriter effect as decoration. It isn’t.

what it looks like the answer is finished on the server and dribbled out to look busy this is the wrong model what is happening the end of the answer does not exist yet when you read the beginning nothing is being withheld There is no “write the whole reply, then send it” mode being hidden from you.
The natural reading is that a finished answer is being revealed slowly for effect. The truth is the opposite — the text is being made as you read it, and every design decision downstream follows from that.

2 · Why the model has no choice

To write word fifty, it has to have already written word forty-nine.

promptprompt + 1prompt + 1 + 2 one pass through the networkone pass through the networkone pass through the network → token 1→ token 2→ token 3 Token N+1 cannot start until token N exists. There is no parallel version of this. a cache keeps each pass from redoing the earlier work — it does not let the passes happen at once
Generation is a loop where each step’s input includes the previous step’s output, which makes it sequential by construction rather than by choice. Streaming isn’t layered on top of a batch process — it is the natural output shape, and buffering is the extra step.

3 · First break: the response wants a length

The ordinary way to send a reply is to announce how big it is. You don’t know.

the usual reply Content-Length: 2417 … which you would have to know before writing a single byte impossible here the open-ended reply chunk · chunk · chunk · … … then an empty one no total announced, ever HTTP has done this since the 1990s No new protocol was needed. The capability was already sitting there. newer HTTP versions frame it differently, but the contract is the same: write bytes now, total unknown
A conventional response declares its size before the body, which a generated answer cannot do. Open-ended bodies were already part of HTTP — the mechanism differs between HTTP versions, but from the application’s side the contract is identical.

4 · Second break: bytes are not messages

What arrives is a hose of bytes, cut in the wrong places.

what the server meant to send messagemessagemessagemessage what actually turns up, read by read half of onethe rest, plus two morea sliverthe remainder So the client buffers, and only parses when it sees a message boundary. a blank line, or a newline, depending on the framing the API chose and decoding must be told the text is unfinished, or a character split across two reads is corrupted
The transport delivers bytes whenever it likes, with no regard for where one logical message ends. The client has to hold an unfinished tail between reads and only parse once a delimiter shows up — the single detail that trips up almost everyone the first time.

5 · Third break: something in the middle helpfully waits

Every “why doesn’t my stream stream?” is a box re-imposing the batch shape.

server chunks, live a proxy, a compressor, a helpful middleware holds it all, to “optimise” delivery one lump, at the end the user, staring The transport almost never breaks streaming. Some box in the middle does. and a client that waits for a complete document before parsing is doing the same thing, on its own side
Proxies, compressors and frameworks that accumulate a whole body before forwarding it are each re-imposing the request-response shape on a process that never had it. Read the bug list that way and it stops being a list — it’s one mistake wearing several costumes.

6 · Keep this card

The whole thing on one index card.

streaming response = an open-ended body, with no length declared + message framing, so bytes become events + flushed the moment the model emits Streaming is the default shape. Buffering is the step someone added. which is also why closing the tab should stop generation — every further token is bought and thrown away
Picture to keep: not a letter that gets posted when it’s finished, but a phone call — the words leave as they’re spoken, and nobody waits for the last sentence before hearing the first.

Why it exists

You ask a chat app to explain something, and the answer types itself out in front of you — a few words, a pause, a few more. You’ve probably read it as a UX flourish, a fancier loading animation. It isn’t. It’s the wire shape of how the model actually produces text, and that one 500-token answer crawling across your screen is the example this post follows.

An LLM has no “write the whole reply, then send it” mode to hide.

It is autoregressive: to produce token N+1 it has to look at tokens 1 through N, including the ones it just produced. There is no parallel “compute the whole reply at once” mode. The fiftieth token is ready after fifty forward passes, in sequence. Each pass takes some milliseconds.

If the server waited for the whole reply before responding, two bad things would happen:

  1. The user stares at a spinner for the full generation time. A 500-token reply at, say, 50 tokens per second is a ten-second wait with nothing on screen, even though the first token was ready well before the rest.
  2. You’d be paying for a buffer you didn’t need. The server holds the text. The network does nothing. The client does nothing. Three idle layers, on purpose.

Streaming is the obvious move once you see this: emit each token (or small group of tokens) as soon as it pops out of the decoder. The user sees motion much sooner than the full reply could possibly arrive, and the same wall-clock generation feels dramatically faster because the time-to-first-token dominates how a chat UI feels — even if it’s not the only latency signal.

Why it matters now

Most major LLM provider APIs ship streaming as a first-class option, and production chat UIs typically default to it. As an engineer in 2026 you’ll hit it from at least three sides:

And it shows up beyond LLMs: agent frameworks stream tool-call deltas and partial structured outputs (some also stream reasoning traces, when the provider exposes them). The pattern is general enough now that “knows how streaming works at the HTTP layer” turns up well outside the teams that build model APIs.

The short answer

streaming response = open-ended HTTP body + message framing + chunks flushed as the model emits them

Picture to keep: not a letter that gets posted when it’s finished, but a phone call — the words leave as they’re spoken, and nobody waits for the last sentence before hearing the first.

Three layers stack here, and confusing them is half the bugs. The model generates incrementally. HTTP carries an open-ended body — via chunked transfer-encoding on HTTP/1.1, via DATA frames on HTTP/2 and HTTP/3. On top of that, an app-level framing — SSE or NDJSON — carves the byte stream into discrete messages the client can parse. The interesting part is what’s causing the chunks: an autoregressive decoder that has no choice but to produce text in order, one step at a time.

How it works

Start from the request-response shape everyone already knows, and watch it break three times.

The decoder is sequential by construction — so buffering wastes the wait

Inside the model, generating a reply is a loop:

prompt → forward pass → token_1
prompt + token_1 → forward pass → token_2
prompt + token_1 + token_2 → forward pass → token_3
...

Each forward pass is one trip through the network’s layers on a GPU. Modern serving stacks reuse most of the work between passes via the KV cache, which is why “tokens per second” is a meaningful steady-state number after the first one. But sequential it remains: token N+1 cannot start until token N exists.

This is the load-bearing fact for the whole post. Streaming isn’t a delivery optimization layered on top of a fundamentally batch computation — it’s the natural output shape of the computation itself. A non-streaming API is the special case, where the server volunteers to buffer.

But an ordinary HTTP response wants to know its length

The naive server writes Content-Length and sends the body — which means knowing how long your 500-token answer is before writing a byte of it, and you don’t. Fix: HTTP already carries open-ended bodies, and you don’t need a new protocol for this. Plain HTTP/1.1 has had chunked transfer-encoding since the 1990s (currently specified in RFC 9112). The body is a sequence of chunks, each prefixed with its length in hex, ending with a zero-length chunk:

HTTP/1.1 200 OK
Content-Type: text/event-stream
Transfer-Encoding: chunked

1a
data: {"token": "Hello"}\n\n
1c
data: {"token": ", world"}\n\n
0

Two common application-level framings sit on top of an open-ended HTTP body:

HTTP/2 and HTTP/3 don’t use Transfer-Encoding: chunked at all — that mechanism is HTTP/1.1-specific. They carry the response body as a sequence of DATA frames on a multiplexed stream, and the server can flush each frame as soon as it’s ready. From the application’s point of view the contract is the same: write some bytes, the other side eventually reads them, and the layer in between doesn’t have to know the total length up front. See QUIC for why the lower layer still matters on flaky links.

But the bytes don’t arrive in message-sized pieces

An open-ended body solves the server’s problem and hands the client a new one: the transport delivers bytes, not events. This is where naïve code falls down. You can’t await response.json() — there is no complete JSON until the server closes the stream. You have to consume the body as it arrives:

const res = await fetch(url, { method: "POST", body: ... });
const reader = res.body.getReader();
const decoder = new TextDecoder();
let buf = "";
while (true) {
  const { value, done } = await reader.read();
  if (done) break;
  buf += decoder.decode(value, { stream: true });
  // split buf on the framing delimiter, parse complete events,
  // keep the unfinished tail for the next iteration.
}

The detail that bites everyone once: a single read() is not aligned to a logical message. You’ll get half an event, then the other half plus two more events, then nothing for 200 ms. The client has to buffer until it sees a complete message boundary (blank line for SSE, newline for NDJSON) before parsing. Decoding bytes-to-text needs the streaming flag too, or you’ll corrupt multi-byte UTF-8 characters whose bytes happened to span a chunk boundary.

Show the seams

A few things the typewriter effect hides:

You started with `streaming response = open-ended HTTP body + message framing

Going deeper