Why synthetic data works for modern LLM training
The open web ran out of high-quality text years before frontier models stopped getting better. The new training signal didn't come from a fresh internet — it came from models writing for models, with filters in front.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- 1. Curated generation from a stronger model
- 2. Distillation as synthetic data
- 3. Self-improvement on verifiable domains
- Why the filter is the magic
- The seam: this is not a free lunch
- Check yourself
- Famous related terms
- Going deeper
The picture version
Five pictures for a reader who assumed the training data all came from people. The prose below fills in the seams the pictures skip.
1 · The problem
The models kept improving. The internet did not get bigger.
2 · The shape of it
Shovel in gravel. Shake the sieve. Keep what’s left.
3 · The obvious objection
A model can’t teach itself something it doesn’t know.
4 · The answer
Every recipe that works has an anchor outside the generator.
5 · Keep this card
The whole thing on one index card.
Why it exists
If you have been paying attention to chatbots since 2023, something doesn’t quite add up. GPT-4 became GPT-4o became GPT-5. Claude 3 became 3.5 became 4.x. Llama 2 became 3 became 4. Every generation got visibly better at coding, at math, at long-form reasoning. But there hasn’t been a new internet to scrape — Reddit and StackOverflow and GitHub didn’t double in size every nine months. So where did the new training signal come from?
The short version is: a lot of it didn’t come from humans at all. The current frontier of LLM training leans heavily on text that other models wrote, filtered hard for the parts that actually teach. This is the thing called synthetic data, and it is one of the loadbearing ingredients of the 2024–2026 generation of models.
The framing comes from a 2022 paper by Villalobos and collaborators at Epoch AI, Will we run out of data? Limits of LLM scaling based on human-generated data (revised 2024). Their projection: at the rates labs were scaling pretraining sets, the stock of high-quality public human text would be exhausted somewhere between roughly 2026 and 2032. That is the data wall. The interesting question is what happened instead of hitting it — because models kept improving, and the wall didn’t visibly stop them.
Why it matters now
If you build on top of frontier models, almost everything you notice about their behavior past the base capability — the way they write code, the way they show their reasoning, the way they format math, the cleanliness of their chain-of-thought — is shaped by post-training on synthetic data. Meta’s The Llama 3 Herd of Models (2024) is the clearest public example: they explicitly describe generating synthetic data for code, math, reasoning, long-context, and tool use, and using it inside SFT and preference-pair generation.
For the closed labs — OpenAI, Anthropic, Google DeepMind — the exact training mixes for GPT-5, Claude 4.x, and Gemini are not public. The shape of the recipe is visible in the open papers; the precise synthetic-to-human ratios at the frontier are not published anywhere, and a specific number would be a guess. What is uncontroversial: synthetic data is no longer a side dish.
The short answer
synthetic data ≈ stronger model writes + filter for what teaches
Picture to keep: a gold-panning tray. The generator shovels in an enormous amount of gravel — cheap, mostly worthless — and the filter is the sieve you shake it through. What survives the sieve is the training set. The useful intuition is that the sieve, not the shovel, is where the value comes from; the analogy breaks in that panning only finds gold that was already in the river, whereas a stronger teacher model really is putting new structure into the gravel before you shake it.
You take a capable existing model (or a verifier, or a chain of both) and have it generate candidate training examples. Then you throw most of them away — keeping only the ones that pass some quality bar. The surviving examples become training data for the next model. The filter is doing as much work as the generator.
How it works
Start with the objection, because it’s the right one. The naive version does not work. Take a model, have it write text, train the same model on that text. You have added no information — you’ve just re-fed the model its own beliefs, sharpened. Do it repeatedly and it gets actively worse (that’s model collapse, below). If synthetic data were only this, the data wall would have stopped the field cold.
Every working recipe is a different answer to the same question: where does the new information enter? The three families below are three answers, and they shade into each other.
1. Curated generation from a stronger model
The cleanest example is Microsoft’s phi line. Textbooks Are All You Need (Gunasekar et al., 2023) trained a 1.3B-parameter code model, phi-1, on a mix of filtered web code and synthetic “textbook-style” Python content generated with GPT-3.5, plus synthetic exercises. Despite being orders of magnitude smaller than contemporaries, phi-1 reached 50.6% pass@1 on HumanEval — the headline result that made “textbook-quality synthetic data” a research direction rather than a curiosity.
The mechanism here is simple: GPT-3.5 already knows how to write a clean, well-commented Python tutorial. The web does not — the web is full of half-finished snippets, copy-pasted answers, and dead StackOverflow threads. Synthetic textbooks let phi-1 learn from concentrated, well-explained code instead of statistical sludge.
The same recipe shows up in Llama 3’s post-training. The Llama 3 paper documents synthetic data pipelines for code (a “code expert” model generating SFT examples), for math, and for reasoning, with quality filters in front and preference-pair generation feeding DPO.
2. Distillation as synthetic data
A teacher model generating outputs that a student trains on is, mechanically, synthetic data — the teacher is the generator. The framing is just different: in distillation you care about compressing a big model into a small one; in synthetic-data work you care about producing examples that teach a behavior. The pipeline is the same. Most of the small open-weights models you have used in the last year were distilled this way, and the line between “distilled” and “trained on synthetic data” is mostly about which side of the conversation you are emphasizing.
3. Self-improvement on verifiable domains
This is the variant with the steepest growth curve, and it deserves its own post — why verifiable domains run away covers it in detail. The setup: in domains where you can automatically check whether an answer is correct (math problems with known answers, code with passing tests, formal proofs), the model can generate millions of attempts, keep only the ones that pass the check, and train on those.
This is why frontier reasoning models have improved on math and code so much faster than on, say, taste in poetry. The verifier is free. Rejection sampling turns “generate” into “generate-and-grade” — the grade is the filter that makes the synthetic data actually useful.
Why the filter is the magic
Now back to the objection we opened with. What all three families have in common is that the filter is not the model that generated the candidates. That’s the whole answer to “where does the new information enter.” The filter is one of:
- a stronger model (teacher → student)
- a verifier (a checker, a unit test, a math grader)
- a rule (length, format, refusal pattern)
- a second model trained to predict human preference (the reward model from RLHF)
In every case there is information entering the training pipeline that didn’t come from the generator. The generator’s job is to produce a wide distribution of candidates; the filter’s job is to extract the parts that are actually correct or useful. The synthetic data is what survives the filter.
This is also why “synthetic data” and “distillation” and “RL on verifiable rewards” are deeply related — they are all the same trick wearing different hats: cheap candidate generation plus a more selective filter.
The seam: this is not a free lunch
If you train a model on its own outputs, and then train the next model on that model’s outputs, and so on, things go wrong. Shumailov et al., The Curse of Recursion: Training on Generated Data Makes Models Forget (2023), showed that recursive training on model-generated data causes model collapse: the tails of the distribution disappear, rare modes get lost, and the model converges to a narrower and narrower slice of what its ancestors knew.
The reason synthetic data works in practice — and doesn’t collapse frontier models — is that the recipe in production isn’t “model trains on its own output, repeat.” It’s:
- A stronger model generates for a weaker one (phi-1 from GPT-3.5; small open models from larger ones). Information flows downhill.
- A verifier filters (only passing solutions survive). The verifier is anchored in something real (a test suite, a known answer), not in the generator.
- Synthetic is mixed with human and web data, not used alone. The Llama 3 paper is explicit about this. The exact ratios aren’t public for the closed labs, but the published recipes all mix.
There are also subtler costs even when collapse is avoided: synthetic data inherits the teacher’s biases, its refusal patterns, its stylistic tics, and its blind spots. If the teacher is wrong in a systematic way, the student inherits the wrongness with high confidence. This is one of the reasons frontier post-training pipelines run many generators, multiple filters, and constant evaluation against held-out human data — to catch the drift before it bakes in.
The standard account is that synthetic data extended the scaling era past where the data wall would have stopped it. Whether it can keep doing so indefinitely — whether the trick has another order of magnitude in it, or whether we’re already in diminishing returns — is genuinely contested in the field.
You started with synthetic data ≈ stronger model writes + filter for what teaches. What did working through the objection add? — + the filter must be anchored outside the generator. That’s the load-bearing clause. A stronger teacher, a unit test, a proof kernel, a human-preference model: each one injects information the generator didn’t have. Drop that anchor and the same pipeline becomes the recursive self-training that Shumailov et al. showed narrows the distribution. Same shovel, no sieve.
Check yourself
Before you go — a small team wants to improve their model’s SQL. They have it generate 100,000 queries, run each against a test database, keep only the queries that execute without error, and fine-tune on those. Will this help? What’s the flaw?
Answer
It will help some, and the flaw is that the verifier is checking the wrong property. “Executes without error” is a real external anchor — that’s genuinely information from outside the generator, and it will scrub away syntax errors and hallucinated column names. But it says nothing about whether the query returns the right rows, so the surviving set is biased toward queries that are safely trivial: SELECT * FROM users LIMIT 10 passes every time. Train on that and you get a model that writes confidently valid, unambitious SQL. The fix is to strengthen the anchor rather than the generator — compare the result set against a labeled expected output, which is exactly the move that turns SQL into a verifiable domain.
And one more — if model collapse is real, why hasn’t it visibly hit frontier models, given how much of their training data is now model-written?
Answer
Because collapse is a result about recursive self-training — a model trained on its own output, repeatedly, with nothing else entering the loop. Production recipes break that loop in at least three places: information flows downhill from a stronger teacher to a weaker student rather than in a circle; verifiers anchored in something real (tests, known answers) do the filtering; and synthetic data is mixed with human and web data rather than replacing it. The honest caveat is that this is an argument about mechanism, not a measurement — the actual synthetic-to-human ratios at frontier labs aren’t public, so “it hasn’t visibly hit them” is an observation about model quality, not a proof that the margin is comfortable.
Famous related terms
- Distillation —
distillation = teacher model + student model + train student on teacher's outputs. Synthetic data with a compression goal. - Rejection sampling —
rejection sampling = generate many + score each + keep the winners. The filter half of synthetic data, especially powerful in verifiable domains. - RLVR —
RLVR = RL loop + reward from a verifier instead of a learned reward model. The training signal is, in effect, filtered synthetic rollouts. - Model collapse —
model collapse ≈ distribution narrowing under recursive self-training. The failure mode synthetic-data pipelines design around. - Data wall —
data wall ≈ projected exhaustion of high-quality public human text. The Villalobos et al. framing the rest of this post sits inside. - Phi series (Microsoft) — small models trained explicitly on filtered + synthetic “textbook-quality” data; the canonical public example of the recipe.
Going deeper
- Villalobos et al., Will we run out of data? Limits of LLM scaling based on human-generated data (2022, revised 2024) — the data-wall framing.
- Gunasekar et al., Textbooks Are All You Need (2023) — phi-1, the cleanest public demonstration that synthetic textbook-style data trains capable code models.
- Meta AI, The Llama 3 Herd of Models (2024) — the most detailed public account of synthetic data inside a frontier post-training pipeline (code, math, reasoning, tool use, long context).
- Shumailov et al., The Curse of Recursion: Training on Generated Data Makes Models Forget (2023) — the model-collapse paper. Read this before assuming synthetic data is free.
- Liu et al., Best Practices and Lessons Learned on Synthetic Data for Language Models (2024, Google DeepMind) — survey-style overview of where synthetic data has worked and where it hasn’t.
Well-established: the data-wall framing, that phi and Llama 3 used synthetic data in the documented ways, that model collapse is real under recursive self-training, and that the filter is doing most of the work. Not public: the synthetic-to-human ratios in current frontier training mixes (GPT-5, Claude 4.x, Gemini). Those mixes have never been published, so any specific ratio you read for them is someone’s estimate.