Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why do scaling laws exist?

Bigger model, more data, more compute — and the loss falls along a straight line on a log-log plot for seven orders of magnitude. Nobody fully knows why that line is so straight.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 16 min read

On this page

The picture version

Six pictures for a reader who has never seen a training budget. The prose below fills in the seams the pictures skip.

1 · The problem

Somebody signs a nine-figure cheque before the thing exists.

pay to the order of one training run $100,000,000 signed — before a single output exists You can’t buy a house that way. You can’t greenlight a film that way. Yet a training run gets budgeted like a bridge. so what does the person signing know that makes it signable?
Before 2020, machine learning had the texture of craft: try an architecture, tune, hope. Nothing about that lets you promise a result in advance — and the cheque is the running example for everything below.

2 · The finding

Plot it right and the points don’t wander. They sit on a straight line.

worse better more parameters, more data, more compute → (each ×10 is one step) every real model lands on the line slide it past the last point Six orders of magnitude, and no bend. That extrapolation is what gets signed against.
Measure how wrong the model is against how much you spent, on logarithmic axes, and the result is a ruler-straight line over a very wide range. Nothing breaks, for a long way — which is what turns a research gamble into a budget.

3 · What changed because of it

The question stopped being “what’s the clever trick?”

“what clever trick wins this benchmark?” research as craft “how much compute buys me this much better?” research as budgeting Which is why model announcements read like logistics reports.
Once error is a known function of spend, the interesting work moves. A large share of frontier effort became operating a known recipe rather than searching for a new one — and the remaining levers are what data you use and what you do after.

4 · The part the line doesn’t settle

Same budget, two ways to spend it — and the first answer was wrong.

one fixed budget, split two ways the first answer a big model not much data later shown to be badly undertrained the revised answer a smaller model far more data — roughly 20 words per parameter And in practice the labs go further still, training long past that. the usual reading: the cheapest model to train isn’t the cheapest to run forever afterwards and the precise ratio is less firmly established than it is usually quoted
The line tells you what a budget buys, not how to divide it between model size and training data — and the first published split was revised sharply. Frontier models then went well past even the revised ratio, because training cost isn’t the only cost that matters.

5 · What the ruler does not measure

The error curve is smooth. What the model can do isn’t.

how wrong it is smooth, predictable, boring whether it can do the task flat for ages, then a jump Some of those jumps are the yardstick, not the model — but either way the two curves differ.
The straight line is about how surprised the model is by text, which is not the same as what it can accomplish. Task performance often sits flat and then jumps — and some of those jumps have been argued to be artefacts of how the task is scored rather than real discontinuities.

6 · Keep this card

The whole thing on one index card.

the law = spend more, and error falls along a straight line — which makes the bill predictable — but says nothing about what it can do and nothing about how to split the budget between model and data Nobody signs on faith. Nobody signs on certainty either. and no settled explanation for why the line is that straight
Picture to keep: a ruler laid across a scatter plot spanning six orders of magnitude — the points sit on its edge, and you can slide it past the last one and read off where the next model lands. Where it breaks, and it matters: it reads off error, not capability, and it says nothing about how far past the last real measurement it stays straight.

Why it exists

You’ve read the headlines: a lab announces it spent hundreds of millions of dollars training one model. Somebody, at some point, signed that check — before the model existed, before anyone had seen a single output. Think about what that requires. You can’t buy a house that way. You can’t greenlight a film that way. Yet a training run gets budgeted like a bridge: this much money, this much time, this much capability at the end. What does anyone know that makes that check signable?

The answer is the thing this post is about. Hold onto that check — it’s the running example for every section below: you have a fixed compute budget, and you have to decide how to spend it.

Most of machine learning before 2020 had the texture of craft. You’d pick an architecture, fiddle with regularization, anneal a learning rate, and hope you’d squeezed another point of accuracy out of the benchmark. There was no reason to believe that bigger models, more data, or more compute would predictably keep paying off — let alone that they’d pay off along a clean mathematical curve. People assumed there was some near-term ceiling. Surely at some point you’d hit diminishing returns; surely the model would start memorizing; surely something would break.

What Kaplan and collaborators at OpenAI showed in early 2020 was that, for transformer language models trained on next-token prediction, nothing breaks. Not for a long, long way. Plot test cross-entropy loss against model size, dataset size, or compute on log-log axes, and you get a straight line — over roughly six orders of magnitude in non-embedding parameters and around eight in adjusted compute. Architecture details (depth vs width, head count) barely matter inside a wide range. What matters is N (parameters), D (data), and C (compute). Inside the measured regime, pick any two and you can predict the loss.

That changed how the field thought about progress. You stopped asking “is there a clever trick that wins this benchmark?” and started asking “how much compute do I need to get loss X?” A lot of frontier-lab work shifted from research into budgeting.

The really uncomfortable part: nobody has a satisfying first-principles explanation for why the line is that straight. There are partial accounts (we’ll get to them), but the empirical fact came first, and a lot of the theory is still catching up.

Why it matters now

Scaling laws are the thing that turned LLM training from research into engineering. They matter to a working engineer for three concrete reasons:

If you’ve ever wondered why every frontier model release reads like an industrial logistics report (“trained on 15T tokens for X exaflop-days”) rather than a research paper, this is why. The recipe is mostly known. The hard part is operating it.

The short answer

scaling law = loss falls as a power law in (N, D, C) — straight line on log-log axes

Picture to keep: a ruler laid across a scatter plot that spans six orders of magnitude — the points don’t wander, they sit on the edge of the ruler, and you can slide the ruler past the last measured point and read off where the next model will land. That extrapolation is what the check gets signed against. Where the ruler analogy breaks, and it matters: it reads off loss, not capability, and it says nothing about how far past the last real data point it stays straight.

If you train transformer language models with the next-token objective and you don’t bottleneck on data or model size, the cross-entropy loss L satisfies, approximately:

L(N) ≈ (Nc / N)^αN     when data isn't the limit
L(D) ≈ (Dc / D)^αD     when model size isn't the limit
L(C) ≈ (Cc / C)^αC     for compute-optimal training

where N is parameter count, D is training tokens, C is compute (flops), and the exponents α are small positive numbers (Kaplan et al. report αN ≈ 0.076 and αD ≈ 0.095 for their setup). The constants Nc, Dc, Cc are not universal — Kaplan et al. are explicit that they depend on the vocabulary and tokenization, so they’re properties of a setup, not of language. The point isn’t the specific numbers — it’s that the relationship is a power law and it holds across many orders of magnitude. (Kaplan et al. 2020.)

How it works

Follow the chain of failures, because that’s how the law actually got built: someone tries the obvious thing, it breaks, the fix becomes the next paragraph.

The naive attempt. You have your compute budget. You do what everyone did before 2020: pick an architecture you like, train it, see what you get, tune, repeat. This fails at the scale of a nine-figure check for a simple reason — it gives you no way to predict the result before spending the money. You need a curve, not a craft.

What was measured

Kaplan et al. trained a swarm of decoder-only transformer language models — varying model size from ~768 to ~1.5B non-embedding parameters, varying data from ~22M to ~23B tokens, varying compute. For each run, they plotted the test cross-entropy loss. Three findings, in plain English:

  1. Loss is a power law in each axis. Plot loss vs N (with enough data and compute), you get a straight line on log-log. Same for D, same for C. No bend.
  2. Architecture details are second-order. Across a wide range of layer counts and aspect ratios, the shape of the model barely changes the curve. The deviations show up at extremes (very few layers, or very lopsided depth/width ratios). What matters is how many non-embedding parameters total.
  3. Sample efficiency improves with size. A bigger model needs proportionally fewer tokens to reach a given loss. This sounds counterintuitive; we’ll come back to it.

Then in 2022, Hoffmann et al. (“Chinchilla”) trained ~400 models from 70M to 16B parameters on 5B–500B tokens, with several methodological differences from Kaplan — including matching the cosine learning-rate schedule’s cycle length to the actual training duration. With those changes, their conclusion shifted: at a fixed compute budget, the compute-optimal mix is equal scaling of model and data — roughly 20 training tokens per parameter, not the much smaller ratios Kaplan’s law implied. They tested it by training a 70B model (“Chinchilla”) on 1.4T tokens, matching Gopher’s compute but using a smaller model with more data, and beat Gopher (280B params, 300B tokens) decisively. Why exactly the two scaling laws disagreed has since been studied carefully; later work (Pearce & Song 2024, Porian et al. 2024) traces the gap to a mix of factors — non-embedding-vs-total parameter conventions, warmup duration, last-layer compute accounting, and optimizer tuning across scales — and argues that learning-rate decay is not the dominant cause.

That’s the headline result. GPT-3 (175B params, 300B tokens — about 1.7 tokens/param) was, by Chinchilla’s lights, dramatically undertrained.

What people argue about

So craft failed and a curve replaced it; then Kaplan’s curve failed on the split between N and D, and Chinchilla replaced that. Does the check sign itself now? Not quite. Two asterisks sit on the Chinchilla story:

Both of these are seams in the textbook account, and worth holding in mind whenever someone cites “Chinchilla-optimal” as if it were a constant of nature.

What nobody fully knows

Here’s the thing that should make you suspicious in a productive way: we have no settled, mechanistic theory for why the loss curve is a power law. There are several proposed explanations and they’re not yet reconciled. Roughly:

None of these is a complete account. The honest summary: we have a robust empirical fact (loss is a power law in N, D, C across many orders of magnitude), several plausible theoretical sketches that each capture part of it, and active disagreement about which sketch is the right one — or whether the right one has been proposed yet.

Where the law breaks

It’s also worth being clear about where the curve isn’t actually flat:

You started with scaling law = loss falls as a power law in (N, D, C). What did following the chain add? — + the split between N and D is a separate, contested decision, and the law is for loss, not for capability. So: the check is signable because the line lets you predict the loss you’ll get for the money. It stays an argument because the line doesn’t tell you what the model can do at that loss, and doesn’t tell you on its own whether to buy parameters or tokens with the budget. That’s the honest answer to the question the headline poses — nobody signs on faith, and nobody signs on certainty either.

Check yourself

Before you go — a lab has a fixed compute budget and two plans: a 70B model on 1.4T tokens, or a 280B model on 300B tokens. Chinchilla says pick the first. Now the lab adds that the model will serve a billion requests a month for two years. Does that change the recommendation, and which quantity did the new fact actually alter?

Answer

It pushes the same direction, but for a reason Chinchilla’s law never mentions. Chinchilla minimizes loss for a fixed training compute budget. “A billion requests a month” adds an inference cost that scales with parameter count and not with training tokens — a 70B model is far cheaper per served token than a 280B one, forever. So what changed is the objective (training cost → training plus lifetime serving cost), and under that objective the optimum slides even further toward “smaller model, more tokens” — past Chinchilla-optimal, which is roughly where Llama 3’s ~15T tokens sits. What did not change is the power law itself. You moved along the same curve; you changed which point on it you wanted.

And one more — a colleague says: “loss curves are smooth, so capabilities improve smoothly; all that talk of models suddenly unlocking abilities is marketing.” Where is that right, and where does it overreach?

Answer

Right on the first half. Cross-entropy loss really is smooth in N, D, and C across the measured range, and Schaeffer et al. 2023 show that many published emergence plots are artifacts of a discontinuous metric (exact-match accuracy, say) sitting on top of smooth underlying improvement — change the metric and the cliff becomes a ramp. It overreaches on the second half. Schaeffer et al. don’t claim every capability jump is an artifact, and “loss is smooth” doesn’t establish that the map from loss to downstream behavior is smooth — that map isn’t itself a power law and isn’t well characterized. The defensible position is narrower than either slogan: smooth loss, contested capability curve.

Going deeper