Why agents fall apart over long horizons
Your agent solves any single step beautifully. Run it for fifty steps and it falls off a cliff. The math behind that cliff is older than LLMs, but a newer twist makes it worse.
On this page
The picture version
Six pictures for a reader who has never built an agent. The prose below fills in the seams the pictures skip.
1 · The problem
Thirty easy steps. It gets the first ten right and then drifts.
2 · The boring half
Being right 95% of the time, thirty times in a row, is being right 21% of the time.
3 · The half that isn’t boring
The risk per step doesn’t stay put. It climbs.
4 · Why the obvious loop feeds it
The simplest agent loop is a bucket that only fills.
5 · What actually helps
None of the fixes make the model smarter.
6 · Keep this card
The whole thing on one index card.
Why it exists
You ask a coding agent — Claude Code, Cursor, Codex, whichever one you have open — to do something boring: add one field to a data model and thread it through the API, the database migration, and the tests. Thirty small steps, none of them hard. This is the running example for the rest of the post. The first ten steps are perfect. Around step twelve it edits the wrong file. At step fifteen it re-reads a doc it read at step four. At step nineteen it “fixes” a test that was already passing. By step twenty-five it is confidently solving a different problem than the one you asked about — and it has never once said it was confused.
The natural reaction is “the model isn’t smart enough yet.” It’s the wrong reaction. You can take a frontier model with near-100% accuracy on a single step and still watch its end-to-end success rate collapse as you stack steps. Sinha et al. (2025), The Illusion of Diminishing Returns, make this concrete: per-step accuracy isn’t constant — it itself degrades as the trajectory gets longer, and even tiny per-step error rates compound multiplicatively across a long chain. A 99% per-step model has a hard time crossing 100 steps without an error somewhere; a 95% model is roughly a coin flip by step 14.
Long-horizon failure isn’t a separate kind of mistake the model makes. It’s a structural property of running a probabilistic text generator in a loop and feeding it its own output.
Why it matters now
This is the bottleneck in current coding, browsing, and computer-use agents.
- Coding agents that pair with you for an hour have to make hundreds of small decisions without a human in the loop between them.
- Research / browsing agents click through dozens of pages and form one answer at the end.
- Computer-use agents — taking screenshots, moving a mouse, typing into an OS — burn through steps fast, and a single confused click can poison the next twenty.
METR’s Measuring AI Ability to Complete Long Tasks (March 2025) put numbers on the trend: the length of tasks frontier agents can complete at 50% reliability has been roughly doubling every ~7 months over the last six years. METR also notes that fitting only the most recent data gives a steeper trend line and that the trend may have accelerated in 2024 — they flag that as uncertainty in the forecast rather than an established result. That’s a real trajectory, but it’s also the flip side of the same fact: the headline capability metric of the era is a length, not an IQ score, because length is exactly where these systems break.
If you’re building agents, the practical consequences are:
- The interesting reliability work is at the harness, not the model. Retries, checkpoints, verifiers, scoped subtasks — that’s what turns a 95% model into a usable system.
- “It works on the demo task” tells you almost nothing. A 5-step scripted demo lives on the easy side of the cliff. Your real workload probably doesn’t.
- Throwing a smarter model at it helps less than you’d expect. The failure is partially structural — the model is conditioning on its own earlier mistakes — and scale doesn’t fully erase that. (More on this below.)
The short answer
long-horizon failure ≈ chain-success arithmetic + a self-conditioning effect that makes per-step accuracy itself degrade as the chain grows
Picture to keep: a chain of paper cups passing water down a line — a little spills at each handoff, so length alone empties the cup; and once a cup is dented, every cup after it is handed a dented cup to copy. Where the picture breaks: a real bucket brigade can’t un-spill, but an agent can be handed a fresh cup — which is exactly what the fixes at the end of this post do.
If each step fails independently with probability p, the chance of
finishing a chain of n steps with no errors is roughly (1−p)ⁿ — that’s
exponential decay in the length of the task even when single-step
accuracy looks great. The newer, more uncomfortable finding is that
per-step accuracy isn’t even constant: once a model’s own context
contains its earlier mistakes, it starts making more of them —
self-conditioning.
The chain isn’t just multiplying a fixed risk; the risk is also rising.
How it works
There are really two things stacked on top of each other. Keeping them separate is the whole point.
1. The boring multiplicative part
Treat your thirty-step field-threading task as a sequence of independent gates and you get the “chain success” identity:
P(finish n steps cleanly) = Π P(step i succeeds)
≈ (1 − p)ⁿ if every step is equally
risky and independent
Before you read the numbers, commit to a guess: your agent is 95% accurate on any single step. Out of 100 attempts at the thirty-step task, how many finish clean?
Plug in numbers and the cliff appears immediately:
| per-step accuracy | 10 steps | 30 steps | 50 steps | 100 steps |
|---|---|---|---|---|
| 99% | 90% | 74% | 60% | 37% |
| 95% | 60% | 21% | 8% | 0.6% |
| 90% | 35% | 4% | 0.5% | ~0% |
So the answer to the guess above: about 21 in 100. A 95% agent fails your boring thirty-step chore roughly four times out of five, and every one of those failures looks like “the model got confused,” never like “the math said so.”
Every percentage point of per-step accuracy buys you a multiplicative win in the length of task you can complete. This is the same arithmetic that has always governed pipelines, manufacturing yield, and any system where you have to nail every step in a row. It just hits harder than people expect because intuitions about accuracy live in the single-question regime.
This part is not specific to LLMs. It’s why “the model is 95% accurate” is almost meaningless without “…over a chain of how long?“.
2. The newer, weirder part: self-conditioning
If per-step error were truly independent and constant, scaling and
fine-tuning would be enough — keep nudging p down and the cliff moves
right. But Sinha et al. (2025) show something more interesting: as a
trajectory grows, the model becomes more likely to make a mistake at
each step, because its context now contains its previous mistakes.
They call this self-conditioning: the model isn’t just sampling
independent errors, it’s sampling conditioned on a transcript that includes
its own bad outputs, and that transcript pulls the next sample further
toward “this is the kind of thing I do.”
A few illustrative shapes this takes in practice (these are intuition, not measured findings from the paper):
- Hallucinated facts get reaffirmed. Once a wrong field name lands in the working notes, the model tends to treat it as evidence and build on it — which is how one bad edit at step twelve becomes a migration and three tests all written against a column that doesn’t exist.
- Bad tool calls beget more bad tool calls. A
grepwith a typo’d pattern returns nothing, which teaches the model — wrongly — that the field isn’t referenced anywhere else, and it starts working from that conclusion. - The agent’s tone shifts. After a few confused steps, you can often see the writing get more apologetic, more tentative, more prone to “let me try a different approach” loops that don’t actually change anything.
The empirical claim from Sinha et al. that’s worth carrying around: this self-conditioning isn’t fully fixed by scaling the model. Bigger models have lower base error rates, but they still degrade across long trajectories from this mechanism. What does help, in their setup, is extended deliberation — they report that recent reasoning models do not show the self-conditioning effect they measure, and report chain-of-thought-prompted variants separately. My read on why — and this is my read, not a result from the paper — is that thinking lets the model re-examine the trajectory before committing the next step, instead of mechanically extending it. That’s a result in their measurement setup; I’m not claiming it generalizes to every agent harness.
Why naive harnesses make this worse
Most “agent loops” in the wild look like:
while not done:
thought, action = model(history)
observation = run(action)
history.append((thought, action, observation))
Two things are quietly hostile to long horizons here:
- The history grows monotonically. Every mistake the model has ever made on this task is in the prompt for every subsequent step. That’s exactly the substrate self-conditioning eats.
- There’s no global state. The “memory” of the agent is the chat log. There’s no separate ledger of “what’s actually true so far,” “which subgoals are done,” “which constraints have been verified.” So every step has to re-derive the world from a transcript that gets noisier over time.
The interventions that work in practice all attack one of these:
- Plan-then-execute. Pin a plan up front — schema, then migration, then API, then tests — treat each step as a scoped subtask with a clear success criterion, and re-plan only when a step actually fails. This bounds how far one bad step can propagate.
- External verifiers / checkers. Run the test suite after the migration edit, run a type-check, diff a file against expectations — anything that turns “the model thinks it added the field” into a machine-checkable signal. Compounding error tolerates wishful thinking; a green check doesn’t.
- Scratchpad pruning / summarization. Periodically replace the long, error-laced history with a clean summary of “where we are” (field added to schema and migration; API and tests outstanding). This is the harness-side answer to self-conditioning: stop feeding the model its own mess.
- Hard reset on detected failure. If a verifier says the step-12 rename broke, rolling back to step 11’s known-good state and retrying is much more reliable than asking the model to “fix it from here.” The state at step 12 is contaminated; the cleanest fix is to throw it out.
- Decompose into shorter horizons. Your thirty-step chore run as three ten-step subtasks chained by a hand-coded glue layer is qualitatively easier than one thirty-step agent run. Note why, because it’s easy to get wrong: splitting alone doesn’t change the arithmetic — three 60% subtasks still multiply out to about 21%. The win comes from the split making each subtask independently checkable and retryable, so a failure costs you ten steps instead of thirty, and from each subtask starting with a clean context instead of inheriting the previous one’s mess.
None of these make the model smarter. They all reduce the number of steps the model has to nail in a row, or break the self-reinforcing-mistake loop. That’s the lever.
Where this argument has limits
A few honest caveats:
- The “independent, equally risky steps” model is a simplification. Real per-step errors are
correlated — some steps are intrinsically harder, some are tutorial-
easy. The clean
(1−p)ⁿcurve is a useful first approximation, not a measurement of any specific agent. - “Self-conditioning” as a name is from a specific 2025 paper. The underlying observation — models doubling down on their own outputs — is older folklore (you’ve seen it in any chatbot that gets stuck in a loop). Sinha et al. give it a tighter operational definition and an experiment; there is no clean comparative result yet against earlier framings of the same phenomenon.
- The METR doubling-time number is for a specific suite of tasks — HCAST, RE-Bench, SWAA, mostly software/research-shaped. How well it extrapolates to other domains (computer-use, embodied agents) is genuinely open. The doubling number is real; the universality is a read, not a measurement.
- A lot of “agent failure” in the wild isn’t pure long-horizon reasoning — it’s bad tool design, bad prompts, bad observability. Multi-agent systems also have their own failure mode taxonomy (Cemri et al., 2025, Why Do Multi-Agent LLM Systems Fail?). The compounding-error argument here is the most common single answer, not the only one.
You started with long-horizon failure ≈ chain-success arithmetic. What did
this post add? — + self-conditioning, and it’s the term that changes what
you build. If the failure were only multiplicative, the entire fix would be
“wait for a better model,” because a lower p moves the cliff right and
nothing else helps. Because the risk also rises inside a run, the fixes
that matter are the ones that shorten the chain or hand the model a clean
cup: verifiers, checkpoints, pruned scratchpads, scoped subtasks.
Which is the answer to the thirty-step chore you opened with. Nothing was wrong with the model at step twelve. The chain was just long enough that something had to go wrong somewhere, and the harness had no way to notice or to throw the bad step away. Length is the dominant axis of difficulty. Smarter models help. Shorter horizons help more.
Check yourself
Before you go — you improve your agent’s per-step accuracy from 95% to 99% on a 50-step task. Roughly how much better does end-to-end success get?
Answer
From about 8% to about 60% — roughly an 8× improvement in success rate, from a 4-point improvement in per-step accuracy. That leverage is the whole reason per-step evals mislead: on a single question the two models look nearly identical (95 vs 99), and on a long chain one is unusable and the other is merely unreliable. It’s also why the last few points of per-step accuracy are worth paying a lot for, and why “our model scores 2% higher on the benchmark” can be a much bigger deal than it sounds.
And a diagnostic: your field-threading agent now reliably derails around step 20 of a 40-step version of the task. You can afford exactly one change — double the context window, or run the test suite after every code edit. Which do you pick, and why?
Answer
The verifier. Unless you were actually truncating, a bigger context window
doesn’t lower p; it just lets you keep more of the trajectory, mistakes
included, which is the substrate self-conditioning feeds on — so it addresses
neither half of the problem.
Running tests after each edit converts “the model believes it added the
field” into a machine-checkable signal, which lets you catch the step-20
error at step 20 instead of at step 40, and lets you roll back to a
known-good state rather than asking a contaminated trajectory to repair
itself. The general rule: buy error detection before you buy error
tolerance.
Famous related terms
- Self-conditioning —
self-conditioning = model sampling errors conditioned on its own past errors in context— Sinha et al.’s name for the “why it gets worse, not just stays bad” effect. - Time horizon (METR) —
time horizon = task duration at which the agent is X% reliable— the metric METR popularized as the measure of agent capability. - ReAct loop —
ReAct = Thought → Action → Observation, repeat— the dominant agent-loop shape; flexible but accumulates history aggressively. - Plan-and-execute —
plan-then-execute = upfront plan + per-step executor + re-planner on failure— limits blast radius of any single bad step at the cost of flexibility. - Reasoning models —
reasoning model = LLM + extended internal deliberation before answering— in Sinha et al.’s setup, dampens self-conditioning. - Agent harness —
agent harness = loop + tools + memory + verifiers around the model— where most reliability work lives. - Hallucination —
hallucination = confidently generated false content— the per-step error type that, once in context, feeds self-conditioning.
Going deeper
- Sinha, Arun, Goel, Staab, Geiping, The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs, arXiv:2509.09677, 2025 — the primary source for self-conditioning, and the place to go for “how would you even measure that per-step accuracy degrades within a run?”
- METR, Measuring AI Ability to Complete Long Tasks, March 2025 — the best explainer of why the field started measuring capability in minutes of task rather than in benchmark points; read the methodology, not the headline number.
- Cemri et al., Why Do Multi-Agent LLM Systems Fail?, arXiv:2503.13657, 2025 — the rabbit hole, for the reader wondering whether adding a second agent fixes any of this (it introduces its own taxonomy of failures instead).