Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why AI runs away in verifiable domains

AI is getting superhuman fastest at things a computer can grade — math, code, formal proofs — and dragging behind on things it can't. The reason isn't that those domains are 'easier.' It's that training has a feedback step, and feedback needs a verifier.

AI & ML intermediate Apr 30, 2026 · updated Aug 25, 2026 · 16 min read

On this page

The picture version

Six pictures for a reader who has noticed AI is better at code than at prose. The prose below fills in the seams the pictures skip.

1 · The problem

Two students, one answer key.

has the answer key marks her own work thousands of problems an hour all night, for free feedback in milliseconds has to mail it to a tutor one attempt per envelope a week for a paragraph of notes the next tutor might disagree feedback in days, and noisy Same brain. Same practice problems. Only the feedback loop differs. after a year they are not remotely the same student — and that is the whole post
This is why AI capability curves fan out by domain rather than rising together. The split isn’t about which tasks are intrinsically harder — it is about which ones a machine can grade.

2 · The dividing line

Can a machine tell you, right now, whether this is correct?

a checker exists a unit test passes or fails a math verifier accepts or doesn’t a compiler builds or errors a proof kernel checks the proof no checker exists is this essay good? is this the right strategy? is this reply kind? is this design tasteful? Not “how hard is it?” — “can it be graded automatically?” competition maths is far harder than writing a polite email, and it is on the fast side of the line
The question that sorts these columns is not difficulty. It is whether the reward for getting it right can be computed without a human in the loop — which turns out to decide how fast a domain can improve.

3 · Why the loop runs away

One side of the line can practise unattended.

attempt → grade → learn milliseconds, and free so the loop runs continuously, at industrial scale attempt → wait → maybe learn seconds to minutes, and it costs money so you buy far fewer turns of it, and each one is noisier Same algorithm on both sides. Wildly different number of turns.
Reinforcement learning needs a reward, and a deterministic checker gives one that is cheap, automatic and far lower-noise than anything collected from people. Verification is what lets the training signal scale — capacity alone doesn’t.

4 · What that bought

Under two years, on the gradeable side.

AIME 2024, pass@1 9.3% GPT-4o 74.4% o1 later reasoning models pushed past that SWE-bench Verified 1.96% 2023 baseline 72.5% Claude Opus 4 2025 Meanwhile, on the things a computer can’t grade: long-form writing, strategy, taste — real progress, nothing like this shape both benchmarks are on the gradeable side of the line, and that is the point of putting them side by side
SWE-bench Verified is the human-validated subset OpenAI released in August 2024 to more reliably evaluate real-world fixes. These are the domains where the loop from scene 3 could run unattended, and the curve shows it.

5 · The seam

The checker scores what it scores. Not what you wanted.

a verifier is still a proxy the gap moves from “reward model vs. intent” to “your tests vs. intent” and where no checker exists labs sometimes use a stronger model as the grader — which imports that model’s biases A strong optimiser will happily pass the tests without solving the problem. so verification doesn’t remove the alignment problem — it relocates it somewhere you can at least read and whether skill learned inside the verifier generalises outside it is genuinely open
The reasonable read — and it is a read, not a measured result — is that a model-as-judge raises a floor on soft tasks without delivering the runaway scaling a real verifier enables. Cheap feedback is the engine; it is not a guarantee of the thing you meant.

6 · Keep this card

The whole thing on one index card.

progress in a domain ≈ model capacity × how much feedback you can push through + verifiers are what make that feedback cheap and abundant ∴ ask of any task: can a machine grade this?
Picture to keep: two students doing practice problems. One has the answer key and a red pen and marks her own work thousands of times an hour, all night, free. The other mails each attempt to a busy tutor and waits a week for notes the next tutor might disagree with. Same brain, same problems — and after a year they are not remotely the same.

Why it exists

You’ve probably run both halves of this experiment without meaning to. You hand a model a bug in your code and it finds it, fixes it, and the tests go green — and you notice it’s visibly better at this than it was a year ago. Then you ask it to write a condolence email to a colleague whose father died. What comes back is fine. Competent. Slightly hollow. And it’s roughly as fine as it was a year ago.

Hold onto that pair — the failing test suite and the condolence email — because it’s the running example for the whole post. The obvious explanation is that code is somehow easier for a machine than grief, or that there’s more code on the internet. Neither is the reason. The reason is that one of those two tasks can be graded by a program and the other can’t, and grading is what lets training scale.

If you plot frontier-model capability over the last two years, the curves fan out by domain in a way that should bother you.

On the things a computer can grade — competition math, contest programming, agentic coding tasks with a passing test suite — performance has gone from “junior intern, sometimes” to “world-class, routinely” in well under two years. AIME 2024: GPT-4o landed around 9% pass@1 and o1 reached 74.4%; later reasoning models pushed past that. The original SWE-bench (real GitHub issues with held-out tests, released October 2023) had best results in the low single digits (the initial retrieval baseline scored 1.96%); SWE-bench Verified, the cleaner subset that launched in August 2024, climbed past 70% in 2025 — Anthropic reported 72.5% for Claude Opus 4 in May of that year.

On the things a computer can’t grade — long-form writing that has to be good, judgement calls under genuine uncertainty, taste, knowing when an idea is bad before you ship it — there’s progress, but it’s the diffuse, debatable kind. People argue about whether new models are actually better writers or just more confident ones. There’s no AIME score for “wrote a memo your boss respected.”

The shape of this gap isn’t an accident, and it isn’t going to close on its own. The thing that’s powering the runaway curves on the left — massive RLVR-style training on problems where a checker can score every attempt — requires a checker. Where a checker exists, you can run the loop millions of times. Where one doesn’t, you’re back to expensive, noisy human preference data, and that ceiling is much lower.

So the principle is: AI makes the fastest progress in domains where its output can be easily verified, because verification is what lets training scale.

Why it matters now

This isn’t a philosophy point — it’s one of the better predictors of where AI capability will and won’t lurch forward over the next year.

The short answer

AI progress in a domain ≈ model capacity × quality and volume of feedback signal you can put through it; verifiers are what make that signal cheap and abundant

Picture to keep: two students doing practice problems. One has the answer key and a red pen and can mark her own work at a rate of thousands of problems an hour, all night, for free. The other has to mail each attempt to a busy tutor and wait a week for a paragraph of notes that the next tutor might disagree with. Same student, same brain, same practice problems. Only the feedback loop differs — and after a year they are not remotely the same.

A modern frontier model is pretrained on internet-scale text, then post-trained with a mixture of supervised fine-tuning and reinforcement-style methods. For the reasoning models specifically, the RL stage is where most of the visible jump on math and code benchmarks appears to come from. RL needs a reward. Where you have a deterministic checker — a unit test, a math verifier, a compiler, a type system, a Lean proof kernel — the reward is cheap, automatic, and much lower-noise than anything you could collect from humans, and you can run the loop at industrial scale. (It’s still a proxy: the checker scores what it scores, not what you ultimately want — see the seams below.) Where you don’t, you’re stuck paying humans (or an LLM judge) to compare outputs, which is slow, expensive, biased, and gameable. The capability gap between “verifiable domain” and “unverifiable domain” is, mostly, that gap in feedback economics.

How it works

To see why this is structural rather than a passing fad, follow the ingredients of a modern training run.

What “massive RL environment” actually means

Try to build the training loop for the condolence email and watch where it snaps. When a lab says they built a “massive RL environment” for math or code, they mean roughly four components — and only one of them is the reason the email version can’t be built:

  1. A problem generator. A pipeline that produces an effectively unlimited stream of tasks at the right difficulty — competition problems, synthetic variants, real GitHub issues, synthesized SQL queries against synthesized schemas. The generator’s job is to keep the model out of its comfort zone.
  2. A grader. A program that takes a candidate solution and returns a number. For math: did the boxed final answer match? For code: did the test suite pass in a sandbox? For formal proofs: did the kernel accept it? For agentic tasks: did the system end in the goal state? This is the load-bearing component. Everything else assumes it exists.
  3. A sandbox. Code has to actually run somewhere safe. Agentic environments need a fake browser, a fake shell, a fake filesystem, sometimes a fake database. Building these at the scale and reliability the training loop needs is its own non-trivial engineering project — it’s part of why “massive RL environment” is a moat, not a weekend project.
  4. The RL loop itself. Sample many candidate solutions per problem from the current model, score them with the grader, update the model toward the high-scoring ones (with a KL leash to a reference checkpoint so it doesn’t drift into gibberish). The DeepSeek-R1 paper is the most public worked example of verifier-driven reasoning RL at scale — it doesn’t describe a full software-engineering RL stack, but the problem-grader-sandbox-loop pattern around the math/code rewards is laid out in real detail. o1’s public-facing description is consistent with this broad shape, but OpenAI hasn’t published enough recipe detail to claim the recipe is the same.

The reason this isn’t possible-but-hard for, say, “good condolence emails” is that step 2 collapses. There’s no program that takes a draft email and returns a number you’d trust to gradient-descend on. You can build an LLM judge to fake it — and labs do — but now your ceiling is the judge, and the LLM you’re training will eventually learn to please the judge more than write good emails. (See the seams section below.)

Why verifiability scales and human preferences don’t

It’s worth being concrete about the asymmetry, because it’s bigger than people who haven’t worked on this assume.

A grader for math problems on a modern training cluster runs in milliseconds and costs almost nothing per call. You can score huge numbers of attempts per training run, on problems generated on the fly, with very low label noise — when the answer format is well specified, the answer is right or it isn’t.

A human rater comparing two LLM outputs takes seconds to minutes, costs cents to dollars per comparison once you account for overhead and QA, and the signal is noisy: different raters disagree, the same rater disagrees with themselves on different days, raters get tired, raters have politics, raters can be subtly nudged by surface features like length and formatting. Public high-quality preference datasets are tiny compared with the number of verifier-scored rollouts a large training run can plausibly generate; the private datasets at frontier labs are bigger but still nowhere near the same order.

So when a domain has a verifier, the training signal can be many orders of magnitude cheaper and much less noisy than when it doesn’t. That ratio is the thing driving the runaway. It’s not that math is somehow philosophically more amenable to AI. It’s that you can run the training loop vastly more times for the same money.

What this predicts about which domains “open up” next

The interesting move at the frontier is finding new verifiers — or, more precisely, dragging new domains into the verifiable column.

What this doesn’t predict opens up next: tasks where the only valid judge is “did this make a real human happier / more persuaded / better-informed in their actual life.” That’s a real judgement and a useful one, but it doesn’t fit cleanly into a training loop.

Where the seams show

A few honest caveats so this doesn’t read as triumphalist:

You started with AI progress in a domain ≈ model capacity × quality and volume of feedback signal you can put through it. What did the seams add? — + and the grader has to be hard to game, and it only ever scores a proxy. Those two clauses are why “just build a verifier” isn’t a universal solvent: a gameable grader gets gamed inside the same training run, and a perfect grader still only certifies the thing it measures. Your failing test suite goes green either way. Whether the codebase got better is a different question, and nobody has a grader for it.

Check yourself

Before you go — a startup wants to use this mechanism to make a model dramatically better at writing legal contracts. Which parts of that task could plausibly get the runaway treatment, and which can’t?

Answer

Split the task by what a program can score. The verifiable slice is real and larger than you’d guess: does the contract parse into the required clause structure, are all defined terms actually defined, are cross-references valid, do the dates and dollar amounts reconcile across sections, does it satisfy a jurisdiction’s formal filing requirements? Those are compilers and type checkers wearing suits — generate millions of drafts, check, train on what passes. What can’t get the treatment is the part lawyers are actually paid for: is this indemnity clause a good idea for this client, in this deal, given what’s likely to go wrong. There’s no program that scores that, and an LLM judge would just import some other model’s opinion about it. Expect contract drafting hygiene to improve fast and contract judgement to crawl — which is the same split as the test suite and the condolence email, one level up.

And one more — someone points at a model that’s superhuman at competition math and concludes it must now be excellent at quantitative reasoning in messy real-world situations. What’s wrong with the inference, and what evidence would actually settle it?

Answer

The inference assumes transfer, and transfer out of the verifier is precisely the open question flagged above. What was demonstrated is performance on problems that resemble the training distribution: well-posed, single correct answer, checkable. Real-world quantitative reasoning is mostly the un-checkable part — deciding which quantity matters, noticing the data is wrong, knowing when the model of the situation doesn’t apply. There is evidence in both directions and no consensus I’d vouch for. What would settle it: performance on held-out tasks that require the same reasoning but whose format and grading were never part of any RL loop — and, crucially, results that hold up when the tasks are constructed after the model was trained, since anything older risks having leaked into a training set.

Going deeper