Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why synthetic data works for modern LLM training

The open web ran out of high-quality text years before frontier models stopped getting better. The new training signal didn't come from a fresh internet — it came from models writing for models, with filters in front.

AI & ML intermediate May 2, 2026 · updated Aug 25, 2026 · 12 min read

On this page

The picture version

Five pictures for a reader who assumed the training data all came from people. The prose below fills in the seams the pictures skip.

1 · The problem

The models kept improving. The internet did not get bigger.

what the models could do the stock of high-quality human text shape of the argument, not plotted data Something has to give — and one published projection puts the crossing between 2026 and 2032.
So where did the extra training data come from? A lot of it didn’t come from humans at all. It came from text other models wrote, filtered hard for the parts that actually teach.

2 · The shape of it

Shovel in gravel. Shake the sieve. Keep what’s left.

a capable model generates candidates enormous, cheap, mostly worthless the filter what survives becomes training data The sieve, not the shovel, is where the value comes from. where the panning analogy breaks: a stronger teacher really is putting new structure into the gravel before you shake it
phi-1 is the cleanest public demonstration: a 1.3B model trained on GPT-3.5-generated textbook-style content reached 50.6% pass@1 on HumanEval. The filter is doing as much work as the generator — which is the claim the next scene tests.

3 · The obvious objection

A model can’t teach itself something it doesn’t know.

generation 1 generation 3 generation 6 Train on your own unfiltered output and the tails disappear. That is real. Shumailov et al. named it model collapse, and the curves above are its shape rather than its measurements so the question isn’t whether the objection is right — it’s what the working recipes do differently
This is the failure mode that makes synthetic data sound like a perpetual-motion machine. It happens specifically under recursive self-training — a generator eating its own output with nothing outside itself to check against.

4 · The answer

Every recipe that works has an anchor outside the generator.

the generator and its student a stronger teacher a test suite, a proof kernel human preference data Each arrow injects information the generator did not already have. drop every arrow and the same pipeline becomes the recursive self-training that collapses. same shovel, no sieve.
And synthetic is mixed with human and web data rather than used alone — the Llama 3 paper is explicit about that. The exact ratios at the frontier have never been published, so any specific number you read for a closed model is someone’s estimate.

5 · Keep this card

The whole thing on one index card.

synthetic data ≈ a stronger model writes + a filter for what teaches ∴ and the filter must be anchored outside the generator
Picture to keep: a gold-panning tray. The generator shovels in an enormous amount of gravel — cheap, mostly worthless — and the filter is the sieve you shake it through. Where the analogy breaks: panning only finds gold already in the river, whereas a stronger teacher really is putting new structure into the gravel first.

Why it exists

If you have been paying attention to chatbots since 2023, something doesn’t quite add up. GPT-4 became GPT-4o became GPT-5. Claude 3 became 3.5 became 4.x. Llama 2 became 3 became 4. Every generation got visibly better at coding, at math, at long-form reasoning. But there hasn’t been a new internet to scrape — Reddit and StackOverflow and GitHub didn’t double in size every nine months. So where did the new training signal come from?

The short version is: a lot of it didn’t come from humans at all. The current frontier of LLM training leans heavily on text that other models wrote, filtered hard for the parts that actually teach. This is the thing called synthetic data, and it is one of the loadbearing ingredients of the 2024–2026 generation of models.

The framing comes from a 2022 paper by Villalobos and collaborators at Epoch AI, Will we run out of data? Limits of LLM scaling based on human-generated data (revised 2024). Their projection: at the rates labs were scaling pretraining sets, the stock of high-quality public human text would be exhausted somewhere between roughly 2026 and 2032. That is the data wall. The interesting question is what happened instead of hitting it — because models kept improving, and the wall didn’t visibly stop them.

Why it matters now

If you build on top of frontier models, almost everything you notice about their behavior past the base capability — the way they write code, the way they show their reasoning, the way they format math, the cleanliness of their chain-of-thought — is shaped by post-training on synthetic data. Meta’s The Llama 3 Herd of Models (2024) is the clearest public example: they explicitly describe generating synthetic data for code, math, reasoning, long-context, and tool use, and using it inside SFT and preference-pair generation.

For the closed labs — OpenAI, Anthropic, Google DeepMind — the exact training mixes for GPT-5, Claude 4.x, and Gemini are not public. The shape of the recipe is visible in the open papers; the precise synthetic-to-human ratios at the frontier are not published anywhere, and a specific number would be a guess. What is uncontroversial: synthetic data is no longer a side dish.

The short answer

synthetic data ≈ stronger model writes + filter for what teaches

Picture to keep: a gold-panning tray. The generator shovels in an enormous amount of gravel — cheap, mostly worthless — and the filter is the sieve you shake it through. What survives the sieve is the training set. The useful intuition is that the sieve, not the shovel, is where the value comes from; the analogy breaks in that panning only finds gold that was already in the river, whereas a stronger teacher model really is putting new structure into the gravel before you shake it.

You take a capable existing model (or a verifier, or a chain of both) and have it generate candidate training examples. Then you throw most of them away — keeping only the ones that pass some quality bar. The surviving examples become training data for the next model. The filter is doing as much work as the generator.

How it works

Start with the objection, because it’s the right one. The naive version does not work. Take a model, have it write text, train the same model on that text. You have added no information — you’ve just re-fed the model its own beliefs, sharpened. Do it repeatedly and it gets actively worse (that’s model collapse, below). If synthetic data were only this, the data wall would have stopped the field cold.

Every working recipe is a different answer to the same question: where does the new information enter? The three families below are three answers, and they shade into each other.

1. Curated generation from a stronger model

The cleanest example is Microsoft’s phi line. Textbooks Are All You Need (Gunasekar et al., 2023) trained a 1.3B-parameter code model, phi-1, on a mix of filtered web code and synthetic “textbook-style” Python content generated with GPT-3.5, plus synthetic exercises. Despite being orders of magnitude smaller than contemporaries, phi-1 reached 50.6% pass@1 on HumanEval — the headline result that made “textbook-quality synthetic data” a research direction rather than a curiosity.

The mechanism here is simple: GPT-3.5 already knows how to write a clean, well-commented Python tutorial. The web does not — the web is full of half-finished snippets, copy-pasted answers, and dead StackOverflow threads. Synthetic textbooks let phi-1 learn from concentrated, well-explained code instead of statistical sludge.

The same recipe shows up in Llama 3’s post-training. The Llama 3 paper documents synthetic data pipelines for code (a “code expert” model generating SFT examples), for math, and for reasoning, with quality filters in front and preference-pair generation feeding DPO.

2. Distillation as synthetic data

A teacher model generating outputs that a student trains on is, mechanically, synthetic data — the teacher is the generator. The framing is just different: in distillation you care about compressing a big model into a small one; in synthetic-data work you care about producing examples that teach a behavior. The pipeline is the same. Most of the small open-weights models you have used in the last year were distilled this way, and the line between “distilled” and “trained on synthetic data” is mostly about which side of the conversation you are emphasizing.

3. Self-improvement on verifiable domains

This is the variant with the steepest growth curve, and it deserves its own post — why verifiable domains run away covers it in detail. The setup: in domains where you can automatically check whether an answer is correct (math problems with known answers, code with passing tests, formal proofs), the model can generate millions of attempts, keep only the ones that pass the check, and train on those.

This is why frontier reasoning models have improved on math and code so much faster than on, say, taste in poetry. The verifier is free. Rejection sampling turns “generate” into “generate-and-grade” — the grade is the filter that makes the synthetic data actually useful.

Why the filter is the magic

Now back to the objection we opened with. What all three families have in common is that the filter is not the model that generated the candidates. That’s the whole answer to “where does the new information enter.” The filter is one of:

In every case there is information entering the training pipeline that didn’t come from the generator. The generator’s job is to produce a wide distribution of candidates; the filter’s job is to extract the parts that are actually correct or useful. The synthetic data is what survives the filter.

This is also why “synthetic data” and “distillation” and “RL on verifiable rewards” are deeply related — they are all the same trick wearing different hats: cheap candidate generation plus a more selective filter.

The seam: this is not a free lunch

If you train a model on its own outputs, and then train the next model on that model’s outputs, and so on, things go wrong. Shumailov et al., The Curse of Recursion: Training on Generated Data Makes Models Forget (2023), showed that recursive training on model-generated data causes model collapse: the tails of the distribution disappear, rare modes get lost, and the model converges to a narrower and narrower slice of what its ancestors knew.

The reason synthetic data works in practice — and doesn’t collapse frontier models — is that the recipe in production isn’t “model trains on its own output, repeat.” It’s:

There are also subtler costs even when collapse is avoided: synthetic data inherits the teacher’s biases, its refusal patterns, its stylistic tics, and its blind spots. If the teacher is wrong in a systematic way, the student inherits the wrongness with high confidence. This is one of the reasons frontier post-training pipelines run many generators, multiple filters, and constant evaluation against held-out human data — to catch the drift before it bakes in.

The standard account is that synthetic data extended the scaling era past where the data wall would have stopped it. Whether it can keep doing so indefinitely — whether the trick has another order of magnitude in it, or whether we’re already in diminishing returns — is genuinely contested in the field.

You started with synthetic data ≈ stronger model writes + filter for what teaches. What did working through the objection add? — + the filter must be anchored outside the generator. That’s the load-bearing clause. A stronger teacher, a unit test, a proof kernel, a human-preference model: each one injects information the generator didn’t have. Drop that anchor and the same pipeline becomes the recursive self-training that Shumailov et al. showed narrows the distribution. Same shovel, no sieve.

Check yourself

Before you go — a small team wants to improve their model’s SQL. They have it generate 100,000 queries, run each against a test database, keep only the queries that execute without error, and fine-tune on those. Will this help? What’s the flaw?

Answer

It will help some, and the flaw is that the verifier is checking the wrong property. “Executes without error” is a real external anchor — that’s genuinely information from outside the generator, and it will scrub away syntax errors and hallucinated column names. But it says nothing about whether the query returns the right rows, so the surviving set is biased toward queries that are safely trivial: SELECT * FROM users LIMIT 10 passes every time. Train on that and you get a model that writes confidently valid, unambitious SQL. The fix is to strengthen the anchor rather than the generator — compare the result set against a labeled expected output, which is exactly the move that turns SQL into a verifiable domain.

And one more — if model collapse is real, why hasn’t it visibly hit frontier models, given how much of their training data is now model-written?

Answer

Because collapse is a result about recursive self-training — a model trained on its own output, repeatedly, with nothing else entering the loop. Production recipes break that loop in at least three places: information flows downhill from a stronger teacher to a weaker student rather than in a circle; verifiers anchored in something real (tests, known answers) do the filtering; and synthetic data is mixed with human and web data rather than replacing it. The honest caveat is that this is an argument about mechanism, not a measurement — the actual synthetic-to-human ratios at frontier labs aren’t public, so “it hasn’t visibly hit them” is an observation about model quality, not a proof that the margin is comfortable.

Going deeper

Well-established: the data-wall framing, that phi and Llama 3 used synthetic data in the documented ways, that model collapse is real under recursive self-training, and that the filter is doing most of the work. Not public: the synthetic-to-human ratios in current frontier training mixes (GPT-5, Claude 4.x, Gemini). Those mixes have never been published, so any specific ratio you read for them is someone’s estimate.