What is an agent harness?
The loop and scaffolding around a language model that turns 'a thing that emits tokens' into 'a thing that does work in the world.'
On this page
The picture version
The whole idea in six pictures, for a reader who has never used a coding agent. The prose below fills in the seams the pictures skip.
1 · The problem
Same model. Wildly different experience.
2 · The naive way
Just ask the model. Text goes in, text comes out.
3 · The fix
Hand it tools, then call it again with what happened.
4 · Where it breaks
Every trip round the loop makes the note longer.
5 · The catch
Notice who is actually holding the keys.
6 · Keep this card
The whole thing on one index card.
Why it exists
You paste an error message — the wall of red text a program spits out when it crashes — into a chat window and ask what’s wrong. You get a genuinely good answer: here are four things to check, in this order. Useful — but you’re the one who has to go open the files, run the commands, and read the output. Now do the same thing in a coding agent: you type “find the slow test and fix it”, and a couple of minutes later there’s a diff on your screen and a test suite that no longer crawls. Same model, wildly different experience. (That’s a composite scenario, not a benchmark result — the point is the shape of the difference, not a specific speedup.)
That’s the gap this post is about, and here’s the surprising part: the model can be identical across those two experiences. A raw LLM maps a sequence of tokens to a probability distribution over the next token. That’s the whole job. It can’t read your files, run your tests, or remember what it did five minutes ago. What changed is everything around the model: a loop, a set of tools it’s allowed to call, some way to track what has happened so far, and a rule for when to stop.
That surrounding machinery is the harness. It exists because the distance between “predicts tokens” and “fixed my slow test” is mostly not model capability — it’s plumbing. The harness is the plumbing. We’ll follow that one request, “find the slow test and fix it,” down through the mechanism.
Why it matters now
Coding agents, IDE copilots, and the tool-using modes of the big chat apps are all harnesses wrapped around a small pool of frontier models — and several of them are wrapped around the same models, since many are built on published provider APIs. The models underneath overlap heavily; the harnesses don’t. Which tools the agent can call, how its context is managed, how it recovers from a failed command, what it’s allowed to do without asking — that’s a large part of where the product actually lives.
It’s also a good first place to look when an agent misbehaves. When one loops forever, edits the wrong file, or forgets an instruction you gave it four steps ago, the first thing to check is whether the right information was ever put in front of the model, and whether the wrong action was ever blocked — both harness decisions. (Some failures are genuinely the model’s. There’s no public breakdown of the split, and I’d distrust anyone who claims a clean number.)
The short answer
agent = model + harness
Picture to keep: a brilliant consultant locked in a windowless room who can only pass notes under the door. The harness is the assistant outside: reading the notes, actually running the commands, sliding the results back in, and deciding which requests are allowed out of the room at all. The analogy breaks in one place worth knowing: the consultant remembers yesterday, and the model doesn’t. Everything it “knows” about this task is whatever the harness put in the current note.
A harness is the ordinary program that runs around a language model — the loop that calls the model, parses what it asks for, executes it, feeds the result back, decides what stays in context, and decides when the task is done. The model supplies judgment; the harness supplies hands, eyes, memory, and a leash.
How it works
Start with the naive version and let it break. Every piece of a real harness is the fix for the previous version’s failure.
Attempt 1: just ask the model. Send "find the slow test and fix it", print
the reply. You get a paragraph of advice. It fails for a boring reason — the
model has never seen your repository, and text is the only thing it can emit.
Fix: tools. Give the model a list of functions it may call, each with a name,
a description, and a
schema
for its arguments — run_bash, read_file,
edit_file. Now instead of prose it can emit a structured request:
{"tool": "run_bash", "command": "pytest --durations=10"}. The harness
validates that request and runs it.
But a single call still gets you nowhere. The model asks to run pytest, the
harness runs it… and the model never sees the timings, because that call already
returned.
Fix: the loop. Append the result to the conversation and call the model again. That’s the entire core of a harness:
context = [system_prompt, user_request]
while True:
response = model(context)
if response.is_final_answer:
return response.text
for tool_call in response.tool_calls:
result = execute(tool_call) # run the bash, read the file…
context.append(tool_call, result)
Our request now goes somewhere:
user: "find the slow test and fix it"
→ model: tool_call(run_bash, "pytest --durations=10")
→ harness runs it, appends the timings
→ model: tool_call(read_file, "tests/test_users.py")
→ harness runs it, appends the file
→ model: tool_call(edit_file, …)
→ harness applies the edit
→ model: "Fixed. The fixture was rebuilding the DB per test."
But the model was never told those tools exist, nor how to behave when a command fails, nor that it shouldn’t rewrite half the repo on the way.
Fix: a system prompt. Persistent instructions prepended to every call — who the agent is, what tools it has, what it should refuse, how to format its answers. In production agents these can run to thousands of tokens; the personality and the policy both live here.
But every loop iteration appends more text, and models have a finite
context window.
A long debugging session — a full pytest log is thousands of tokens on its
own — overflows it, and then something has to be dropped. Drop the wrong thing
and the agent forgets the instruction you gave it at the start.
Fix: context management. The harness decides what stays verbatim, what gets summarized, what gets dropped, and what can be re-fetched on demand. It’s the part that’s easiest to get wrong, and it’s why an agent that was sharp for ten minutes goes vague at minute forty.
But now the model can emit anything, and the harness will run it. Our
agent, asked to speed up a test suite, is one plausible-looking tool call away
from rm -rf on a directory it misread.
Fix: permissions. Before a destructive call executes, the harness pauses and asks you, or refuses outright. Note where the authority sits: the model can only propose. The harness decides what actually runs. Approval prompts are annoying for exactly the reason they’re load-bearing.
But the loop as written has no exit except the model volunteering one. An agent that can’t find the slow test will keep trying forever, on your budget.
Fix: a stopping condition. The model emits a done-signal, or the user interrupts, or a budget in tokens/time/cost is exhausted, or a guardrail trips.
Which explains the thing that seems paradoxical from outside: the same weights in two different harnesses behave like two different products, and a good harness around a weaker model can beat a sloppy harness around a stronger one. “This agent feels smart” and “this model is smart” are separate claims.
An honest caveat about the ordering above: this is a re-derivation, not a history. Tool-calling loops, context management, and permission systems were developed across many groups roughly in parallel, not discovered in this sequence.
You started with agent = model + harness. After walking the loop, what’s the
one word that best describes what the harness contributes? — authority. The
harness holds the tools, the memory, and the off switch; the model only ever
gets to ask.
Famous related terms
- Agent —
agent = model + harness— the whole running thing, model plus the scaffolding that lets it act. - Tool use / function calling —
tool use = model emits structured calls + harness executes them— the protocol that turns “describe how to run pytest” into actually running it. - Context window —
context window = how many tokens the model can see at once— the fixed budget every context-management decision is fighting over. - System prompt —
system prompt = persistent instructions prepended to every model call— where an agent’s personality and its refusal policy both live. - ReAct —
ReAct ≈ reason + act, interleaved— the early pattern of alternating between thinking out loud and emitting a tool call. - MCP (Model Context Protocol) —
MCP = open protocol for exposing tools, resources, and prompts to AI apps— a shared protocol, so a harness that speaks MCP can pick up a new tool server without bespoke integration code for each one. - Scaffolding —
scaffolding ≈ harness— another word for the same idea; you’ll see both, often in the same conversation.
Going deeper
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022) — the paper to read if you want the original argument for why interleaving reasoning with tool calls beats doing either alone.
- Anthropic, Building effective agents — answers “when do I actually need a loop, versus a fixed workflow that’s cheaper and more predictable?”
- The OpenHands source (or Aider’s, or Continue’s) — for the reader who wants to find out how much of a real harness turns out to be unglamorous parsing, retries, and prompt assembly.