Why dropout disappeared from modern LLMs
Dropout was the regularization workhorse of the deep-learning era. Frontier LLM pretraining quietly stopped using it. The reason isn't that dropout broke — it's that the problem dropout solved stopped being the problem.
On this page
The picture version
Five pictures for a reader who has never tuned a regularizer. The prose below fills in the seams the pictures skip.
1 · Two students
One re-reads a single book. One reads fifty, once each.
2 · What dropout does
Hide a random third of the network on every single pass.
3 · What changed
The threat dropout was built for stopped showing up.
4 · The regime check
Fine-tune on 4,000 tickets and you are student one again.
5 · Keep this card
The whole thing on one index card.
Why it exists
Picture two ways to study for an exam. The first: re-read the same textbook five times until you’ve practically memorized the page numbers. The second: read fifty different textbooks once each. The first student crushes practice questions taken from that textbook and then bombs the real exam, because they learned the book, not the subject. The second student has never seen any single passage twice — and ends up understanding the material better. Dropout was invented for the first kind of student. Modern LLM pretraining looks like the second kind, and that’s most of the story.
In the 2010s, neural networks were trained on relatively small, repeated datasets, and they would happily overfit — fit the training set perfectly while failing on anything new. Srivastava, Hinton and coauthors published Dropout: A Simple Way to Prevent Neural Networks from Overfitting in JMLR in 2014, and the trick was as memorable as the name: during training, randomly zero out a fraction of activations on every forward pass. The network can’t lean on any single neuron, so it learns more distributed, robust features. It worked, it was cheap to implement, and it became the default — the canonical CNN and BERT-era transformer recipes shipped with dropout on. The original Attention Is All You Need (Vaswani et al., 2017) used Pdrop = 0.1 in its base model (and 0.3 in one of the big variants); BERT (Devlin et al., 2018) used 0.1 in pretraining and fine-tuning.
Then the data and the models got much bigger, and the recipes quietly stopped turning it on.
Why it matters now
This isn’t just historical trivia — it’s a working example of how scaling changes which knobs matter:
- Recipes copied from BERT-era code mostly still set
dropout=0.1. If you fine-tune or pretrain a small model with that default, you’re paying a regularization tax that may not buy you anything at your scale. The defaults outlived the regime they were designed for. - Fine-tuning is a different regime from pretraining. When you fine-tune on a small task-specific dataset, you are the 2015 student re-reading one textbook. Dropout (and its cousins like attention dropout) often comes back on for fine-tuning and LoRA adapters even when pretraining ran without it.
- The mental model “more regularization = better generalization” stops carrying its weight at scale. When you’ve spent a Chinchilla-optimal compute budget, anything that adds noise to gradients without buying matching generalization eats into effective steps. Whether dropout specifically interacts badly with fused attention kernels or low-precision activations has no clean public source — the more defensible point is just that the upside shrank as data scale grew, while the engineering surface area didn’t.
The short answer
dropout-free pretraining = scale + data diversity > stochastic regularization
Picture to keep: dropout is the tutor who hides a random third of the textbook’s pages each time the first student reads it, forcing them to reconstruct rather than recite. Do that to the second student — the one already reading fifty different books once each — and you’re just tearing pages out of books they were only going to see once anyway. The analogy breaks in one place: dropout isn’t hiding data, it’s zeroing the network’s own internal activations. But the shape of the bargain is the same — noise you add to prevent memorization is only worth it if memorization was the threat.
Dropout fights overfitting by injecting noise so the network can’t lean on individual features. At internet-scale pretraining, the model sees most tokens roughly once, and the data distribution is broad enough that the repeated-example overfitting pressure dropout was designed to fight is largely absent. (LLMs still memorize — verbatim recall of training data is a real, measured phenomenon — but that’s a different failure mode than the one dropout addresses.) Data scale and diversity do most of the regularization work for free, and the per-step noise dropout adds stops paying its keep.
How it works
To see why scale changes the equation, it helps to remember what dropout actually was doing.
Dropout sets a random subset of activations to zero on each training step (typically 10–50% in the 2010s). The standard story has two parts. (1) It approximates an exponentially large ensemble: each forward pass is a different sub-network, and at inference you “average” them by using the full network with rescaled activations. (2) It prevents co-adaptation: no neuron can rely on any specific other neuron being present, so features have to stand on their own.
Both stories assume the model is in a regime where the same training examples are seen many times and the network can latch onto incidental co-adaptations. That regime used to be the default — a CNN trained on ImageNet sees each image dozens of times across epochs. It isn’t anymore for frontier pretraining.
Frontier LLM pretraining looks different in three ways that matter:
- Single-pass-ish data. Pretraining runs are typically one epoch or close to it over trillions of tokens. LLaMA 1’s data mix, for example, uses most components for one epoch, with Wikipedia and books at roughly two. The model rarely sees the same exact sequence many times.
- The dataset itself does much of the regularizing. Web text, code, books, math, multilingual data — the distribution is broad enough that any feature the model learns has to pay rent across many domains. This is the field’s working mental model, not a proven mechanism, but it lines up with what open recipes converged on.
- Compute, not variance, is the binding constraint. Once you’re allocating tokens and parameters under a fixed compute budget (the Chinchilla framing), anything that injects noise into gradients without buying matching generalization eats into effective steps. At small scale, dropout’s noise pays for itself in better generalization. At large scale with diverse data, that bargain shifts.
What do open-weight technical reports actually say? The LLaMA 1 paper (Touvron et al., 2023) does not report using dropout in pretraining. Pythia’s released config sets attention and hidden dropout to 0 for pretraining. Other open-weight reports in the same era describe similar setups — low or zero dropout for pretraining, sometimes nonzero for downstream fine-tuning. The exact configurations of closed frontier models (GPT-4, Claude, Gemini) are not public, so I can’t tell you their dropout rates. What’s public is the trend in open-weight reports and the underlying argument: when the data is doing the regularizing, the noise injection isn’t earning its slot.
The honest seam: the closest published ablation is Liu, Bauer and Manning’s Drop Dropout on Single Epoch Language Model Pretraining (Findings of ACL 2025), which studies dropout in single-epoch pretraining at BERT-base and Pythia 160M–1.4B scale and reports that performance improves when dropout is switched off during pretraining. But that isn’t 70B+ frontier scale, and no public head-to-head holds tokens, parameters, optimizer, and precision fixed at frontier scale — the labs that could run it don’t publish training configs. The case in this post is therefore partly mechanistic (the data-diversity argument), partly observational (open-weight recipes converged on low/zero dropout), and partly path-dependent. The safe claim is: dropout’s role shrank dramatically as pretraining scale grew, not that it has been formally proven harmful.
You started with dropout-free pretraining = scale + data diversity > stochastic regularization. What did this post add about why dropout is still shipping in half the code you’ll read? — + a regime check. Dropout didn’t stop working; the two students never stopped existing. Fine-tuning on your 5,000-row dataset is still the first student re-reading one textbook, and that’s exactly where you should expect the knob to earn its keep again.
Check yourself
Before you go — you’re fine-tuning an open 8B model on 4,000 labelled support tickets. The pretraining config you inherited has dropout=0.0. Keep it, or turn it on?
Answer
Probably turn it on — or reach for some other regularizer. The argument in this post is about regime, not about dropout being obsolete: pretraining sees trillions of diverse tokens roughly once, so the repeated-example overfitting pressure is largely absent. Four thousand tickets over several epochs is the opposite situation — a narrow distribution seen many times, which is precisely the setting dropout was designed for. The inherited config isn’t wrong; it was tuned for a different problem. (In practice you’d sweep it rather than trust either default, and LoRA adapters ship their own dropout parameter for exactly this reason.)
And one more, trickier: someone argues that since LLMs demonstrably memorize training data verbatim, they must be overfitting, so dropout should help. Where does that argument go wrong?
Answer
It conflates two different phenomena that both get called memorization. Classical overfitting is training loss falling while held-out loss rises — the model fitting noise in a small repeated dataset at the cost of generalizing. Verbatim recall in an LLM is the model storing specific sequences while held-out loss keeps improving; it’s a capacity-and-duplication effect, not a generalization failure of the kind dropout targets. The evidence for it is real, but it isn’t evidence that the pretraining run is in the regime dropout was built for. Fixing it looks like deduplicating the corpus, not adding activation noise.
Famous related terms
- Weight decay —
weight decay = loss + λ·||weights||²— the other classic regularizer. Unlike dropout, it survived the scale transition and is standard in modern LLM training. It penalizes weight magnitude rather than injecting activation noise, which is friendlier to large-batch optimization. - Label smoothing —
label smoothing = one-hot target + small uniform mass— softens the training target. Used in the original transformer; usage in modern LLM pretraining is mixed and not always disclosed. - Data augmentation —
data augmentation = training set + label-preserving transformations— the vision-world cousin of “more data fixes it.” LLMs effectively get this for free from the diversity of web text. - Why scaling laws exist —
scaling laws ≈ loss falls predictably as parameters, data, and compute grow— the broader story for why “make it bigger and feed it more data” reshaped which tricks matter. - Why fine-tuning is cheap —
fine-tuning ≈ pretraining minus building representations from scratch— the regime where dropout often comes back on, because the data is small and overfitting is real again.
Going deeper
- Srivastava, Hinton et al. — Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR 2014) — the primary source, for “what did dropout claim to do, and on what kind of dataset was that claim tested?” The answer to the second half is most of this post.
- Hoffmann et al. — Training Compute-Optimal Large Language Models (Chinchilla, 2022) — the explainer for the regime shift, if your question is “why did ‘more tokens per parameter’ become the default, and what does that do to a fixed step budget?”
- Liu, Bauer, Manning — Drop Dropout on Single Epoch Language Model Pretraining (Findings of ACL 2025) — the rabbit hole, and the one place someone actually ran the ablation: what happens to BERT-base and Pythia 160M–1.4B when you switch dropout off for a single-epoch pretraining run.
- Touvron et al. — LLaMA (2023) — read the training-details section and notice what isn’t there. Comparing open-weight recipes for what they omit is a useful habit in general.