Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

What is an agent harness?

The loop and scaffolding around a language model that turns 'a thing that emits tokens' into 'a thing that does work in the world.'

AI & ML intro Apr 29, 2026 · updated Aug 24, 2026 · 9 min read

On this page

The picture version

The whole idea in six pictures, for a reader who has never used a coding agent. The prose below fills in the seams the pictures skip.

1 · The problem

Same model. Wildly different experience.

THE SAME MODEL in a chat window in a coding agent “here are four things to check, in this order” your job now you open the files, run the commands, and read the output a diff on your screen, and a test suite that no longer crawls a couple of minutes you typed one sentence: “find the slow test and fix it”
The model can be identical across these two experiences. What changed is everything around it — and that surrounding machinery is the harness.

2 · The naive way

Just ask the model. Text goes in, text comes out.

THE MODEL a windowless room can't read your files can't run your tests can't remember five minutes ago the door slot “check whether a fixture is rebuilding the DB” a note — the only thing it can emit your repo files, tests, a terminal nothing crosses this gap on its own
It fails for a boring reason: the model has never seen your repository, and text is the only thing it can emit. You get advice, and you do the work.

3 · The fix

Hand it tools, then call it again with what happened.

MODEL picks the next thing worth trying HARNESS validates it, runs it, appends the result a structured request, not prose {"tool": "run_bash", "command": "pytest"} the output, glued onto the conversation … and around again, until the model says it's done you: “find the slow test and fix it” → run_bash  pytest --durations=10 → read_file tests/test_users.py → edit_file … ✓ “Fixed. The fixture was rebuilding the DB per test.” each arrow is one trip around the loop
That loop is the entire core of a harness: call the model, run what it asks for, append the result, call it again. One call alone gets you nowhere — the model would never see the timings.

4 · Where it breaks

Every trip round the loop makes the note longer.

… and this doesn't fit more log, more files, more turns system prompt — your standing rules “find the slow test and fix it” the full pytest log thousands of tokens on its own tests/test_users.py the context window — a fixed height, always what survives? THE HARNESS CHOOSES keep it word for word summarise it drop it fetch it again if needed drop the wrong thing → it forgets your instruction
This is why an agent that was sharp for ten minutes goes vague at minute forty. Nothing is remembered — things are only re-sent, and the harness decides what gets re-sent.

5 · The catch

Notice who is actually holding the keys.

MODEL can only propose rm -rf … proposes HARNESS decides pauses and asks you, or refuses outright and holds the off switch approved actually runs never happens blocked the loop ends when the harness says so: a done-signal, your interrupt, a spent budget, or a tripped guardrail
The model only ever gets to ask. Approval prompts are annoying for exactly the reason they're load-bearing — a plausible-looking tool call is one step away from a directory the model misread.

6 · Keep this card

The whole thing on one index card.

agent = model + harness model brings judgment harness brings hands, eyes, memory, and a leash same weights, different harness → different product
Picture to keep: a brilliant consultant locked in a windowless room who can only pass notes under the door — and the assistant outside who runs the commands, slides the results back in, and decides which requests are allowed out of the room at all.

Why it exists

You paste an error message — the wall of red text a program spits out when it crashes — into a chat window and ask what’s wrong. You get a genuinely good answer: here are four things to check, in this order. Useful — but you’re the one who has to go open the files, run the commands, and read the output. Now do the same thing in a coding agent: you type “find the slow test and fix it”, and a couple of minutes later there’s a diff on your screen and a test suite that no longer crawls. Same model, wildly different experience. (That’s a composite scenario, not a benchmark result — the point is the shape of the difference, not a specific speedup.)

That’s the gap this post is about, and here’s the surprising part: the model can be identical across those two experiences. A raw LLM maps a sequence of tokens to a probability distribution over the next token. That’s the whole job. It can’t read your files, run your tests, or remember what it did five minutes ago. What changed is everything around the model: a loop, a set of tools it’s allowed to call, some way to track what has happened so far, and a rule for when to stop.

That surrounding machinery is the harness. It exists because the distance between “predicts tokens” and “fixed my slow test” is mostly not model capability — it’s plumbing. The harness is the plumbing. We’ll follow that one request, “find the slow test and fix it,” down through the mechanism.

Why it matters now

Coding agents, IDE copilots, and the tool-using modes of the big chat apps are all harnesses wrapped around a small pool of frontier models — and several of them are wrapped around the same models, since many are built on published provider APIs. The models underneath overlap heavily; the harnesses don’t. Which tools the agent can call, how its context is managed, how it recovers from a failed command, what it’s allowed to do without asking — that’s a large part of where the product actually lives.

It’s also a good first place to look when an agent misbehaves. When one loops forever, edits the wrong file, or forgets an instruction you gave it four steps ago, the first thing to check is whether the right information was ever put in front of the model, and whether the wrong action was ever blocked — both harness decisions. (Some failures are genuinely the model’s. There’s no public breakdown of the split, and I’d distrust anyone who claims a clean number.)

The short answer

agent = model + harness

Picture to keep: a brilliant consultant locked in a windowless room who can only pass notes under the door. The harness is the assistant outside: reading the notes, actually running the commands, sliding the results back in, and deciding which requests are allowed out of the room at all. The analogy breaks in one place worth knowing: the consultant remembers yesterday, and the model doesn’t. Everything it “knows” about this task is whatever the harness put in the current note.

A harness is the ordinary program that runs around a language model — the loop that calls the model, parses what it asks for, executes it, feeds the result back, decides what stays in context, and decides when the task is done. The model supplies judgment; the harness supplies hands, eyes, memory, and a leash.

How it works

Start with the naive version and let it break. Every piece of a real harness is the fix for the previous version’s failure.

Attempt 1: just ask the model. Send "find the slow test and fix it", print the reply. You get a paragraph of advice. It fails for a boring reason — the model has never seen your repository, and text is the only thing it can emit.

Fix: tools. Give the model a list of functions it may call, each with a name, a description, and a schema for its arguments — run_bash, read_file, edit_file. Now instead of prose it can emit a structured request: {"tool": "run_bash", "command": "pytest --durations=10"}. The harness validates that request and runs it.

But a single call still gets you nowhere. The model asks to run pytest, the harness runs it… and the model never sees the timings, because that call already returned.

Fix: the loop. Append the result to the conversation and call the model again. That’s the entire core of a harness:

context = [system_prompt, user_request]
while True:
    response = model(context)
    if response.is_final_answer:
        return response.text
    for tool_call in response.tool_calls:
        result = execute(tool_call)        # run the bash, read the file…
        context.append(tool_call, result)

Our request now goes somewhere:

user: "find the slow test and fix it"

→ model: tool_call(run_bash, "pytest --durations=10")
→ harness runs it, appends the timings
→ model: tool_call(read_file, "tests/test_users.py")
→ harness runs it, appends the file
→ model: tool_call(edit_file, …)
→ harness applies the edit
→ model: "Fixed. The fixture was rebuilding the DB per test."

But the model was never told those tools exist, nor how to behave when a command fails, nor that it shouldn’t rewrite half the repo on the way.

Fix: a system prompt. Persistent instructions prepended to every call — who the agent is, what tools it has, what it should refuse, how to format its answers. In production agents these can run to thousands of tokens; the personality and the policy both live here.

But every loop iteration appends more text, and models have a finite context window. A long debugging session — a full pytest log is thousands of tokens on its own — overflows it, and then something has to be dropped. Drop the wrong thing and the agent forgets the instruction you gave it at the start.

Fix: context management. The harness decides what stays verbatim, what gets summarized, what gets dropped, and what can be re-fetched on demand. It’s the part that’s easiest to get wrong, and it’s why an agent that was sharp for ten minutes goes vague at minute forty.

But now the model can emit anything, and the harness will run it. Our agent, asked to speed up a test suite, is one plausible-looking tool call away from rm -rf on a directory it misread.

Fix: permissions. Before a destructive call executes, the harness pauses and asks you, or refuses outright. Note where the authority sits: the model can only propose. The harness decides what actually runs. Approval prompts are annoying for exactly the reason they’re load-bearing.

But the loop as written has no exit except the model volunteering one. An agent that can’t find the slow test will keep trying forever, on your budget.

Fix: a stopping condition. The model emits a done-signal, or the user interrupts, or a budget in tokens/time/cost is exhausted, or a guardrail trips.

Which explains the thing that seems paradoxical from outside: the same weights in two different harnesses behave like two different products, and a good harness around a weaker model can beat a sloppy harness around a stronger one. “This agent feels smart” and “this model is smart” are separate claims.

An honest caveat about the ordering above: this is a re-derivation, not a history. Tool-calling loops, context management, and permission systems were developed across many groups roughly in parallel, not discovered in this sequence.

You started with agent = model + harness. After walking the loop, what’s the one word that best describes what the harness contributes? — authority. The harness holds the tools, the memory, and the off switch; the model only ever gets to ask.

Going deeper