Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why agents fall apart over long horizons

Your agent solves any single step beautifully. Run it for fifty steps and it falls off a cliff. The math behind that cliff is older than LLMs, but a newer twist makes it worse.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 14 min read

On this page

The picture version

Six pictures for a reader who has never built an agent. The prose below fills in the seams the pictures skip.

1 · The problem

Thirty easy steps. It gets the first ten right and then drifts.

“add one field, and thread it through the code” steps 1–10: perfect step 12: edits the wrong file step 19: “fixes” a passing test step 25: a different problem And it never once says it is confused.
None of the thirty steps is hard, and a good model gets almost any single one right. The failure appears only when you stack them — which is the first clue that it isn’t really about how smart the model is.

2 · The boring half

Being right 95% of the time, thirty times in a row, is being right 21% of the time.

100 attempts at the thirty-step chore 21 clean runs. 79 confusing failures. what the chain does to a good model 100% 0 steps in the chain → 99% per step still loses most 100-step tasks. this part is arithmetic, not artificial intelligence
Success on a chain is every step’s success multiplied together, so it falls off a cliff as the chain grows. A 95%-per-step agent fails your thirty-step chore about four times in five — and each failure looks like confusion rather than like arithmetic.

3 · The half that isn’t boring

The risk per step doesn’t stay put. It climbs.

what you’d assume same chance of a mistake every step what actually happens the chance of a mistake grows as it goes Every step reads the transcript of every step before it — mistakes included. So the mistakes become the example to follow.
Because the agent’s only memory is its own transcript, a bad step at twelve is still sitting in front of it at step twenty as if it were evidence. That pushes the per-step error rate up as the run gets longer — researchers call it self-conditioning, and a bigger model does not remove it.

4 · Why the obvious loop feeds it

The simplest agent loop is a bucket that only fills.

the model decides one step does the step whatever happened gets appended the history the bad step, still here nothing is ever removed, corrected, or checked There is no separate record of what is actually true so far. so every step has to re-derive the world from a transcript that keeps getting noisier
The plain loop appends thought, action and result forever and calls that memory. That growing transcript is exactly the material the rising error rate feeds on — and there is no ledger anywhere saying which parts of it were ever verified.

5 · What actually helps

None of the fixes make the model smarter.

check it, don’t take its word run the tests, run the type-checker, diff the file compounding error tolerates wishful thinking; a green check doesn’t throw the bad step away roll back to the last state that passed, rather than asking a confused run to repair itself the contaminated state is the problem, not the model hand it a clean page replace the long, error-laced history with a short summary of where things actually stand stop feeding it its own mess make the chains shorter three checkable ten-step jobs, not one thirty-step job splitting alone changes nothing — the win is that a failure now costs ten steps, and can be retried Every one of them either shortens the chain or breaks the feedback loop.
The lever is never “think harder.” It is fewer steps that have to be nailed in a row, and a way to notice and discard a bad one — splitting a task helps because each piece becomes independently checkable and retryable, not because the multiplication changes.

6 · Keep this card

The whole thing on one index card.

long-horizon failure = every step’s odds, multiplied + a risk that rises inside the run because the run keeps reading its own mistakes the second term is why “wait for a better model” is not the whole answer
Picture to keep: a line of paper cups passing water — a little spills at every handoff, and once a cup is dented every cup after it is handed a dented cup to copy. Unlike a real bucket brigade, an agent can be handed a fresh cup, and that is what every fix above does. Length, not difficulty, is the axis that breaks these systems.

Why it exists

You ask a coding agent — Claude Code, Cursor, Codex, whichever one you have open — to do something boring: add one field to a data model and thread it through the API, the database migration, and the tests. Thirty small steps, none of them hard. This is the running example for the rest of the post. The first ten steps are perfect. Around step twelve it edits the wrong file. At step fifteen it re-reads a doc it read at step four. At step nineteen it “fixes” a test that was already passing. By step twenty-five it is confidently solving a different problem than the one you asked about — and it has never once said it was confused.

The natural reaction is “the model isn’t smart enough yet.” It’s the wrong reaction. You can take a frontier model with near-100% accuracy on a single step and still watch its end-to-end success rate collapse as you stack steps. Sinha et al. (2025), The Illusion of Diminishing Returns, make this concrete: per-step accuracy isn’t constant — it itself degrades as the trajectory gets longer, and even tiny per-step error rates compound multiplicatively across a long chain. A 99% per-step model has a hard time crossing 100 steps without an error somewhere; a 95% model is roughly a coin flip by step 14.

Long-horizon failure isn’t a separate kind of mistake the model makes. It’s a structural property of running a probabilistic text generator in a loop and feeding it its own output.

Why it matters now

This is the bottleneck in current coding, browsing, and computer-use agents.

METR’s Measuring AI Ability to Complete Long Tasks (March 2025) put numbers on the trend: the length of tasks frontier agents can complete at 50% reliability has been roughly doubling every ~7 months over the last six years. METR also notes that fitting only the most recent data gives a steeper trend line and that the trend may have accelerated in 2024 — they flag that as uncertainty in the forecast rather than an established result. That’s a real trajectory, but it’s also the flip side of the same fact: the headline capability metric of the era is a length, not an IQ score, because length is exactly where these systems break.

If you’re building agents, the practical consequences are:

The short answer

long-horizon failure ≈ chain-success arithmetic + a self-conditioning effect that makes per-step accuracy itself degrade as the chain grows

Picture to keep: a chain of paper cups passing water down a line — a little spills at each handoff, so length alone empties the cup; and once a cup is dented, every cup after it is handed a dented cup to copy. Where the picture breaks: a real bucket brigade can’t un-spill, but an agent can be handed a fresh cup — which is exactly what the fixes at the end of this post do.

If each step fails independently with probability p, the chance of finishing a chain of n steps with no errors is roughly (1−p)ⁿ — that’s exponential decay in the length of the task even when single-step accuracy looks great. The newer, more uncomfortable finding is that per-step accuracy isn’t even constant: once a model’s own context contains its earlier mistakes, it starts making more of them — self-conditioning. The chain isn’t just multiplying a fixed risk; the risk is also rising.

How it works

There are really two things stacked on top of each other. Keeping them separate is the whole point.

1. The boring multiplicative part

Treat your thirty-step field-threading task as a sequence of independent gates and you get the “chain success” identity:

P(finish n steps cleanly) = Π P(step i succeeds)
                          ≈ (1 − p)ⁿ      if every step is equally
                                          risky and independent

Before you read the numbers, commit to a guess: your agent is 95% accurate on any single step. Out of 100 attempts at the thirty-step task, how many finish clean?

Plug in numbers and the cliff appears immediately:

per-step accuracy10 steps30 steps50 steps100 steps
99%90%74%60%37%
95%60%21%8%0.6%
90%35%4%0.5%~0%

So the answer to the guess above: about 21 in 100. A 95% agent fails your boring thirty-step chore roughly four times out of five, and every one of those failures looks like “the model got confused,” never like “the math said so.”

Every percentage point of per-step accuracy buys you a multiplicative win in the length of task you can complete. This is the same arithmetic that has always governed pipelines, manufacturing yield, and any system where you have to nail every step in a row. It just hits harder than people expect because intuitions about accuracy live in the single-question regime.

This part is not specific to LLMs. It’s why “the model is 95% accurate” is almost meaningless without “…over a chain of how long?“.

2. The newer, weirder part: self-conditioning

If per-step error were truly independent and constant, scaling and fine-tuning would be enough — keep nudging p down and the cliff moves right. But Sinha et al. (2025) show something more interesting: as a trajectory grows, the model becomes more likely to make a mistake at each step, because its context now contains its previous mistakes. They call this self-conditioning: the model isn’t just sampling independent errors, it’s sampling conditioned on a transcript that includes its own bad outputs, and that transcript pulls the next sample further toward “this is the kind of thing I do.”

A few illustrative shapes this takes in practice (these are intuition, not measured findings from the paper):

The empirical claim from Sinha et al. that’s worth carrying around: this self-conditioning isn’t fully fixed by scaling the model. Bigger models have lower base error rates, but they still degrade across long trajectories from this mechanism. What does help, in their setup, is extended deliberation — they report that recent reasoning models do not show the self-conditioning effect they measure, and report chain-of-thought-prompted variants separately. My read on why — and this is my read, not a result from the paper — is that thinking lets the model re-examine the trajectory before committing the next step, instead of mechanically extending it. That’s a result in their measurement setup; I’m not claiming it generalizes to every agent harness.

Why naive harnesses make this worse

Most “agent loops” in the wild look like:

while not done:
    thought, action = model(history)
    observation     = run(action)
    history.append((thought, action, observation))

Two things are quietly hostile to long horizons here:

The interventions that work in practice all attack one of these:

None of these make the model smarter. They all reduce the number of steps the model has to nail in a row, or break the self-reinforcing-mistake loop. That’s the lever.

Where this argument has limits

A few honest caveats:

You started with long-horizon failure ≈ chain-success arithmetic. What did this post add? — + self-conditioning, and it’s the term that changes what you build. If the failure were only multiplicative, the entire fix would be “wait for a better model,” because a lower p moves the cliff right and nothing else helps. Because the risk also rises inside a run, the fixes that matter are the ones that shorten the chain or hand the model a clean cup: verifiers, checkpoints, pruned scratchpads, scoped subtasks.

Which is the answer to the thirty-step chore you opened with. Nothing was wrong with the model at step twelve. The chain was just long enough that something had to go wrong somewhere, and the harness had no way to notice or to throw the bad step away. Length is the dominant axis of difficulty. Smarter models help. Shorter horizons help more.

Check yourself

Before you go — you improve your agent’s per-step accuracy from 95% to 99% on a 50-step task. Roughly how much better does end-to-end success get?

Answer

From about 8% to about 60% — roughly an 8× improvement in success rate, from a 4-point improvement in per-step accuracy. That leverage is the whole reason per-step evals mislead: on a single question the two models look nearly identical (95 vs 99), and on a long chain one is unusable and the other is merely unreliable. It’s also why the last few points of per-step accuracy are worth paying a lot for, and why “our model scores 2% higher on the benchmark” can be a much bigger deal than it sounds.

And a diagnostic: your field-threading agent now reliably derails around step 20 of a 40-step version of the task. You can afford exactly one change — double the context window, or run the test suite after every code edit. Which do you pick, and why?

Answer

The verifier. Unless you were actually truncating, a bigger context window doesn’t lower p; it just lets you keep more of the trajectory, mistakes included, which is the substrate self-conditioning feeds on — so it addresses neither half of the problem. Running tests after each edit converts “the model believes it added the field” into a machine-checkable signal, which lets you catch the step-20 error at step 20 instead of at step 40, and lets you roll back to a known-good state rather than asking a contaminated trajectory to repair itself. The general rule: buy error detection before you buy error tolerance.

Going deeper