Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why image generation went diffusion, not autoregressive

LLMs are autoregressive: predict the next token. Image models could have been the same — predict the next pixel. Almost none of the dominant ones are. Here's why the field walked away from that approach.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 17 min read

On this page

The picture version

Six pictures for a reader who has only ever watched the progress bar. The prose below fills in the seams the pictures skip.

1 · The problem

One types at you. The other resolves out of static.

ask a chatbot for a paragraph The fox sat next? one word after the last — you can start reading before it finishes ask an image generator for a fox pure static a blurry shape a fox not one corner at a time. the whole picture, sharpening at once. two fundamentally different machines
Not a user-interface choice. The two systems factorise the problem in incompatible ways, and everything downstream — cost per image, why there is no “streaming first pixels”, why one sampling step looks blurry — falls out of that.

2 · The road not taken

“Predict the next pixel” is a real design. It loses three ways.

shaded = not chosen yet an order you had to invent 1 · the ordering is a fiction pixel (100,100) depends on (101,101) exactly as much as the reverse — the factorisation has to pick a side anyway 2 · early commitments cannot be revisited by the time it picks the chin, the eyes and hairline are already fixed. small errors bake in and drift. 3 · the sequence is brutal 512 × 512 RGB, pixel by pixel ≈ 786,000 steps learned image tokenizers cut this down sharply — Parti uses a 32×32 = 1,024-token grid for a 256×256 image — but a thousand sequential transformer steps with quadratic attention is still an expensive way to make one picture MaskGIT and friends sit in the same token family without being raster-order AR; the branch that lost is specifically “next image token, in a fixed order”
Autoregression solves a problem images don’t have — a natural ordering — while ignoring two they do have. Text gets away with it because text really is roughly causally ordered; pixels aren’t.

3 · The move

Train on a task so dull it barely looks generative.

TRAINING — repeated over millions of (image, timestep) pairs a real image add noise at step t a noisy image ONE NETWORK shown the noisy image and the number t its guess at the noise that was added the only thing it ever learns loss = squared error between guess and actual noise No adversary. No discriminator. No pixel ordering. one regression problem — which is exactly why it stayed stable when the field scaled it up
That is the entire training objective. Notice what is not in it: nothing about generating anything. The generative model is not the network — it is what happens when you run the network many times in a row.

4 · The payoff

Generating is just running the schedule backwards.

−noise −noise −noise step T step 0 and every arrow is the same network, called again no ordering the network sees the entire canvas every time and may change any part of it steps are a knob 1000 down to 20–50, and 1–4 for distilled models — a compute dial, not a property of the image errors are recoverable a region that came out wrong at step 30 is just part of the input at step 29 One procedure appends. The other rewrites.
All three columns are consequences of the same refusal to pick an order. Nothing is ever committed — every step takes the whole current canvas as input and hands back a whole new one, which is precisely what an autoregressive sample cannot do.

5 · The practical version

Run the whole thing somewhere much smaller.

512×512×3 pixel space — expensive encode 64×64×4 the whole denoising loop happens in here 64×64×4 ×N steps a much smaller denoiser, run many times — the compute per step drops by orders of magnitude decode back to pixels — lossily This is what made diffusion cheap enough to open-source. the autoencoder is not lossless — it is lossy in ways the diffusion stage doesn’t care about, which is the trade the latent-diffusion paper makes explicit
Stable Diffusion up to SDXL is approximately “latent diffusion + a text encoder fed in via cross-attention + a public release.” The lineage has since split — SD3-class models moved to flow-matching internals, so “diffusion” is now a product label covering several distinct pieces of maths.

6 · Keep this card

The whole thing on one index card.

diffusion model = a denoiser + a fixed noise schedule + sample by reversing it from pure noise ∴ no ordering — which is the whole answer to scene 1
Picture to keep: a photograph developing in a darkroom tray, run in reverse and repeatedly — you start with pure grain and each dip pulls a little more image out of the noise. Where it breaks, and it matters: the darkroom reveals something already latent on the paper, while the denoiser is inventing what should have been there. That is why two runs from different noise give two different foxes.

Why it exists

You’ve watched both of these happen. Ask a chatbot for a paragraph and it types at you, left to right, one word appearing after the last — you can start reading before it finishes. Ask an image generator for a picture of a fox and you get something completely different: a rectangle of static that gradually resolves into a blurry shape, then a fox. Not one corner at a time. The whole picture, sharpening at once. That contrast is the running example for this post, and it isn’t a UI choice — it’s two fundamentally different machines.

If you came to generative AI through LLMs, the first machine sounds like the whole story. Giant transformer, predict the next token, sample from it. Text, code, JSON — same recipe.

So the obvious question is: why isn’t the fox the same trick? Pixels are just numbers. You could flatten an image into a sequence of pixel values and predict the next one. Some early models did exactly that — PixelRNN and PixelCNN in 2016, Image GPT in 2020. They worked. They just didn’t win. For most of the last five years the dominant open-weights and publicly documented image/video systems — DDPM, Stable Diffusion (1.x/2.x/SDXL), Imagen, and Sora on the video side — have been diffusion models (or close cousins like flow matching). Different machine, different objective, different sampling procedure. For closed systems like Midjourney, Runway, and Veo the exact internals aren’t public, so treat “diffusion won” as a claim about the cluster of papers and open weights, not a per-product attribution.

This post is about why. The short version is that “predict the next pixel” is a real architecture, but it solves a problem images don’t have (a natural ordering) while ignoring problems they do have (every pixel matters at once, and tiny independent errors compound). Diffusion gave up on sequence prediction entirely and replaced it with something that fits the geometry of images much better: start from noise, denoise toward an image, in many small steps.

Engineers integrating image or video generation into products keep running into the consequences — cost per image, latency, why a single sampling step looks blurry, why guidance scale matters, why there’s no “streaming first pixels” the way there’s streaming first tokens. Those are all downstream of this choice.

Why it matters now

Diffusion is the production form factor for visual generation:

If you’re shipping anything visual, the cost model, the failure modes, and the controllability primitives all come from this design choice.

The short answer

diffusion model = a denoiser + a fixed noise schedule + sample by reversing the schedule from pure noise

Picture to keep: a photograph developing in a darkroom tray, except run in reverse and repeatedly — you start with pure grain and each dip pulls a little more of the image out of the noise. The analogy breaks where it matters most: the darkroom is revealing something that was already latent on the paper, whereas the denoiser is inventing what should have been there, guessing at every step. That’s why two runs from different noise give two different foxes.

You train one neural network to do one job: given a noisy image and a number telling it how noisy, predict the noise (or equivalently, a slightly-cleaner version). To generate a new image, you start with pure random noise and apply that denoiser many times in sequence, each step removing a little more noise, until what’s left is an image. The “generative model” is the entire reverse-noising trajectory, not a single forward pass.

That’s it. The rest of the post is why this beat autoregressive pixels for images, and where the seams are.

How it works

To see why diffusion fits images, it helps to first see what’s wrong with the LLM-style approach when you point it at pixels.

The problem with “predict the next pixel”

An autoregressive image model has to choose an order. Top-left to bottom-right? Hilbert curve? Coarse-to-fine over patches? Whatever you pick, you’re now claiming that pixel (i, j) only depends on pixels that came earlier in your ordering. That’s not how images work. Pixel (100, 100) depends on pixel (101, 101) just as much as the other way around. The autoregressive factorization picks a side anyway, because it has to.

This bites in three ways:

  1. Long-range coherence is hard. By the time the model is choosing a pixel near the bottom of a face, it has already committed to the pixels near the top — eyes, hairline. If the bottom doesn’t match (chin shape, lighting), the model can’t go back. LLMs have the same problem in principle, but text is roughly causally ordered (we read left-to-right, the next word does mostly depend on prior words). Pixels aren’t.
  2. Errors compound multiplicatively. Every sampled pixel conditions on every previous sampled pixel. A small mistake early — a slightly-off skin tone — gets baked into the conditional distribution for everything after. With millions of pixels per image, the joint distribution drifts.
  3. Sequence length is brutal. A 512×512 RGB image is ~786,000 “tokens” if you go pixel-by-pixel. Image tokenizers cut this down sharply — Parti, for example, uses a 32×32 = 1,024-token grid for a 256×256 image — but the per-token cost of an autoregressive transformer plus quadratic attention still makes naive AR image generation expensive enough that early pixel-RNN/CNN models could only ever produce small images.

You can fix some of this. Image GPT attacked the length problem crudely — downsample the image and shrink the colour palette, then predict that shorter pixel sequence — while later token-based image models (Parti, plus the autoregressive part of GPT-4o image generation) use far more sophisticated learned image tokenizers. There are also non-autoregressive token-based generators in the same family: MaskGIT (Chang et al., CVPR 2022) deliberately isn’t raster-scan AR; it predicts masked tokens in parallel and iteratively refines, which is closer in spirit to diffusion than to next-token prediction. The “predict the next image token in raster order” branch is the one that lost, not “all token-based image models.”

What diffusion actually does

Diffusion flips the problem. Instead of generating an image one piece at a time, it generates the whole image at every step, but starts with one that’s almost entirely noise and gradually removes the noise.

The training story is small enough to hold in your head:

  1. Take a real image x.
  2. Pick a random “timestep” t between 0 (clean) and T (pure noise).
  3. Add Gaussian noise to x according to a fixed schedule that knows how much noise corresponds to step t. Call the result x_t.
  4. Show the network x_t and t. Ask it to predict the noise that was added (this is the DDPM noise-prediction objective from Ho, Jain, Abbeel 2020).
  5. Loss = mean squared error between predicted and actual noise.
  6. Repeat over millions of (image, timestep) pairs.

That’s the entire training objective. No adversarial loss, no likelihood-by-pixel-ordering, no discriminator. One regression problem.

To generate, you reverse the schedule:

  1. Sample pure noise x_T.
  2. For t = T, T-1, ..., 1: ask the network “what noise is in x_t at step t?” Subtract a fraction of it. You now have x_{t-1}, slightly less noisy.
  3. After T steps you’ve reached x_0, an image.

A few things fall out of this that are worth pausing on:

Latent diffusion: the practical version

The DDPM paper (Ho, Jain, Abbeel 2020) ran diffusion in pixel space. That works for small images, but for 512×512 or 1024×1024 it’s expensive — you’re running a U-Net over the full image at every step.

Latent diffusion (Rombach et al., CVPR 2022; this is the architecture behind Stable Diffusion) added one trick: train an autoencoder first that compresses images into a smaller latent space (e.g. 512×512 RGB becomes a 64×64×4 tensor), and then run diffusion in that latent space. The denoiser is smaller, the per-step compute is much lower, and a (lossy) autoencoder handles the high-frequency detail. The LDM paper frames this explicitly as a perceptual-quality vs. compression trade — the autoencoder is not lossless, just lossy in ways the diffusion stage doesn’t care about.

The open-weights image diffusion lineage since — Stable Diffusion 1.x, 2.x, SDXL, and plenty of the third-party fine-tunes — is built on this idea. “Stable Diffusion” up to SDXL is approximately “latent diffusion + a text encoder fed in via cross-attention + a public release.” SD3 moved to flow-matching internals, so the lineage is no longer a single recipe.

Why the field bet on this

The pivotal empirical moment was Dhariwal and Nichol, Diffusion Models Beat GANs on Image Synthesis (2021). Up to that point GANs led the headline image-quality benchmarks. Dhariwal and Nichol showed that on ImageNet at 256×256 and 512×512, diffusion beat them on FID, and it was more stable to train besides. Note the scope: that’s a claim about those benchmarks, not a proof that diffusion dominates every image task. GANs are notoriously fiddly — mode collapse, training instability, hyperparameter sensitivity. The diffusion training loop is boring in comparison: one regression objective, no discriminator, no adversarial dynamics. That mattered a lot when scaling up.

So what diffusion gave the field, in plain terms:

The cost is that you need many forward passes per sample. That’s a real, painful trade — and a huge amount of recent work (consistency models, rectified flow, distillation) is about pushing the step count down without tanking quality.

Where diffusion misleads you

You started with `diffusion model = a denoiser + a fixed noise schedule

Check yourself

Before you go — a colleague argues that since diffusion needs 20–50 forward passes and an LLM needs one pass per token, diffusion must be the more expensive design. Under what conditions is that wrong?

Answer

It’s wrong whenever the autoregressive alternative would need more than 20–50 passes, which for images is essentially always. A pixel-by-pixel model of a 512×512 image needs hundreds of thousands of sequential steps; even with a good image tokenizer you’re at roughly a thousand. Diffusion’s step count is a knob you choose, decoupled from the size of the output — which is why it’s roughly fixed latency per image, while autoregressive cost scales with how much output there is. The comparison your colleague is making would be fair between two diffusion models, not between diffusion and AR pixels.

And one more: the post says a bad diffusion step can be partially undone by later steps, whereas a bad autoregressive token can’t. What is it about the two sampling procedures that makes that true?

Answer

An autoregressive sample is committed: once a token is drawn it becomes part of the condition for everything after, and there’s no operation in the procedure that revisits it. A diffusion step doesn’t commit to anything final — every step takes the entire current canvas as input and outputs a new entire canvas, so a region that came out wrong at step 30 is just part of the input at step 29 and can be pushed back toward the data distribution. The asymmetry is structural: one procedure appends, the other rewrites.

Going deeper

Well-established: the noise-prediction training objective, the reasons autoregressive pixels lost (ordering, error compounding, sequence length), and the latent-diffusion compute argument. Not public: the share of commercial frontier systems still running “classical” diffusion vs. flow matching vs. autoregressive image tokens. Labs publish selectively, and marketing labels don’t always match the math underneath — so the load-bearing claim here is about the cluster of published work, not about any individual product.