Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why dropout disappeared from modern LLMs

Dropout was the regularization workhorse of the deep-learning era. Frontier LLM pretraining quietly stopped using it. The reason isn't that dropout broke — it's that the problem dropout solved stopped being the problem.

AI & ML intermediate May 2, 2026 · updated Aug 25, 2026 · 10 min read

On this page

The picture version

Five pictures for a reader who has never tuned a regularizer. The prose below fills in the seams the pictures skip.

1 · Two students

One re-reads a single book. One reads fifty, once each.

student one the book read it five times learns the book, not the subject — then bombs the real exam student two read fifty, once each never sees a passage twice — and understands the material better Dropout was invented for the student on the left. modern pretraining looks like the student on the right, and that is most of the story
The failure dropout was built to prevent is overfitting: fitting a small, repeated training set perfectly and then failing on anything new. Hold on to which student you are — the whole post is a check on that.

2 · What dropout does

Hide a random third of the network on every single pass.

pass 1 pass 2 pass 3 filled = switched off this pass a different sub-network every single training step no unit can rely on any other so features have to stand on their own and it approximates a huge ensemble averaged at inference by running the full network both stories assume the same examples come round again and again — which is the assumption that stopped holding
Cheap to implement, memorable, and it worked: the canonical CNN and BERT-era transformer recipes shipped with it on. The original Attention Is All You Need used 0.1 in its base model; BERT used 0.1 in pretraining and fine-tuning.

3 · What changed

The threat dropout was built for stopped showing up.

the 2010s regime one image set seen dozens of times round and round frontier pretraining web text code books maths multilingual trillions of tokens, most of them seen roughly once The data is already doing the regularizing. any feature the model learns has to pay rent across many domains — the field’s working mental model, not a proven mechanism and under a fixed compute budget, noise that doesn’t buy matching generalization is just eating into effective steps
LLaMA 1’s mix uses most components for a single epoch, with Wikipedia and books at roughly two; its paper reports no pretraining dropout, and Pythia’s released config sets attention and hidden dropout to zero. Dropout didn’t break — the problem it solved stopped being the problem.

4 · The regime check

Fine-tune on 4,000 tickets and you are student one again.

how many times will it see the same example? ask this, not “is dropout good?” roughly once many times pretraining trillions of varied tokens, one pass the data regularizes. the noise stops paying its keep. fine-tuning 4,000 support tickets, several epochs a narrow distribution, seen again and again. turn it back on. Dropout didn’t stop working. The two students never stopped existing. which is why the inherited dropout=0.0 in a pretraining config isn’t wrong — it was tuned for a different problem
The argument in this post is about regime, not about dropout being obsolete. The honest limit: the closest published ablation runs at BERT-base and Pythia 160M–1.4B scale, not 70B+, and no public head-to-head holds tokens, parameters, optimizer and precision fixed at frontier scale.

5 · Keep this card

The whole thing on one index card.

dropout-free pretraining = scale + data diversity > stochastic regularization + a regime check ∴ noise beats memorization only if memorization was the threat
Picture to keep: dropout is the tutor who hides a random third of the textbook’s pages each time the first student reads it, forcing them to reconstruct rather than recite. Do that to the second student and you are just tearing pages out of books they were only going to see once anyway.

Why it exists

Picture two ways to study for an exam. The first: re-read the same textbook five times until you’ve practically memorized the page numbers. The second: read fifty different textbooks once each. The first student crushes practice questions taken from that textbook and then bombs the real exam, because they learned the book, not the subject. The second student has never seen any single passage twice — and ends up understanding the material better. Dropout was invented for the first kind of student. Modern LLM pretraining looks like the second kind, and that’s most of the story.

In the 2010s, neural networks were trained on relatively small, repeated datasets, and they would happily overfit — fit the training set perfectly while failing on anything new. Srivastava, Hinton and coauthors published Dropout: A Simple Way to Prevent Neural Networks from Overfitting in JMLR in 2014, and the trick was as memorable as the name: during training, randomly zero out a fraction of activations on every forward pass. The network can’t lean on any single neuron, so it learns more distributed, robust features. It worked, it was cheap to implement, and it became the default — the canonical CNN and BERT-era transformer recipes shipped with dropout on. The original Attention Is All You Need (Vaswani et al., 2017) used Pdrop = 0.1 in its base model (and 0.3 in one of the big variants); BERT (Devlin et al., 2018) used 0.1 in pretraining and fine-tuning.

Then the data and the models got much bigger, and the recipes quietly stopped turning it on.

Why it matters now

This isn’t just historical trivia — it’s a working example of how scaling changes which knobs matter:

The short answer

dropout-free pretraining = scale + data diversity > stochastic regularization

Picture to keep: dropout is the tutor who hides a random third of the textbook’s pages each time the first student reads it, forcing them to reconstruct rather than recite. Do that to the second student — the one already reading fifty different books once each — and you’re just tearing pages out of books they were only going to see once anyway. The analogy breaks in one place: dropout isn’t hiding data, it’s zeroing the network’s own internal activations. But the shape of the bargain is the same — noise you add to prevent memorization is only worth it if memorization was the threat.

Dropout fights overfitting by injecting noise so the network can’t lean on individual features. At internet-scale pretraining, the model sees most tokens roughly once, and the data distribution is broad enough that the repeated-example overfitting pressure dropout was designed to fight is largely absent. (LLMs still memorize — verbatim recall of training data is a real, measured phenomenon — but that’s a different failure mode than the one dropout addresses.) Data scale and diversity do most of the regularization work for free, and the per-step noise dropout adds stops paying its keep.

How it works

To see why scale changes the equation, it helps to remember what dropout actually was doing.

Dropout sets a random subset of activations to zero on each training step (typically 10–50% in the 2010s). The standard story has two parts. (1) It approximates an exponentially large ensemble: each forward pass is a different sub-network, and at inference you “average” them by using the full network with rescaled activations. (2) It prevents co-adaptation: no neuron can rely on any specific other neuron being present, so features have to stand on their own.

Both stories assume the model is in a regime where the same training examples are seen many times and the network can latch onto incidental co-adaptations. That regime used to be the default — a CNN trained on ImageNet sees each image dozens of times across epochs. It isn’t anymore for frontier pretraining.

Frontier LLM pretraining looks different in three ways that matter:

What do open-weight technical reports actually say? The LLaMA 1 paper (Touvron et al., 2023) does not report using dropout in pretraining. Pythia’s released config sets attention and hidden dropout to 0 for pretraining. Other open-weight reports in the same era describe similar setups — low or zero dropout for pretraining, sometimes nonzero for downstream fine-tuning. The exact configurations of closed frontier models (GPT-4, Claude, Gemini) are not public, so I can’t tell you their dropout rates. What’s public is the trend in open-weight reports and the underlying argument: when the data is doing the regularizing, the noise injection isn’t earning its slot.

The honest seam: the closest published ablation is Liu, Bauer and Manning’s Drop Dropout on Single Epoch Language Model Pretraining (Findings of ACL 2025), which studies dropout in single-epoch pretraining at BERT-base and Pythia 160M–1.4B scale and reports that performance improves when dropout is switched off during pretraining. But that isn’t 70B+ frontier scale, and no public head-to-head holds tokens, parameters, optimizer, and precision fixed at frontier scale — the labs that could run it don’t publish training configs. The case in this post is therefore partly mechanistic (the data-diversity argument), partly observational (open-weight recipes converged on low/zero dropout), and partly path-dependent. The safe claim is: dropout’s role shrank dramatically as pretraining scale grew, not that it has been formally proven harmful.

You started with dropout-free pretraining = scale + data diversity > stochastic regularization. What did this post add about why dropout is still shipping in half the code you’ll read? — + a regime check. Dropout didn’t stop working; the two students never stopped existing. Fine-tuning on your 5,000-row dataset is still the first student re-reading one textbook, and that’s exactly where you should expect the knob to earn its keep again.

Check yourself

Before you go — you’re fine-tuning an open 8B model on 4,000 labelled support tickets. The pretraining config you inherited has dropout=0.0. Keep it, or turn it on?

Answer

Probably turn it on — or reach for some other regularizer. The argument in this post is about regime, not about dropout being obsolete: pretraining sees trillions of diverse tokens roughly once, so the repeated-example overfitting pressure is largely absent. Four thousand tickets over several epochs is the opposite situation — a narrow distribution seen many times, which is precisely the setting dropout was designed for. The inherited config isn’t wrong; it was tuned for a different problem. (In practice you’d sweep it rather than trust either default, and LoRA adapters ship their own dropout parameter for exactly this reason.)

And one more, trickier: someone argues that since LLMs demonstrably memorize training data verbatim, they must be overfitting, so dropout should help. Where does that argument go wrong?

Answer

It conflates two different phenomena that both get called memorization. Classical overfitting is training loss falling while held-out loss rises — the model fitting noise in a small repeated dataset at the cost of generalizing. Verbatim recall in an LLM is the model storing specific sequences while held-out loss keeps improving; it’s a capacity-and-duplication effect, not a generalization failure of the kind dropout targets. The evidence for it is real, but it isn’t evidence that the pretraining run is in the regime dropout was built for. Fixing it looks like deduplicating the corpus, not adding activation noise.

Going deeper