What is harness engineering?
Most of the work that turns a frontier model into a reliable product happens around the model, not inside it. Harness engineering is the name for that work.
On this page
- The picture version
- Why it exists
- Why it matters now
- The short answer
- How it works
- Break 1: it has no hands → tool design
- Break 2: thirty steps in, it forgets the job → context engineering
- Break 3: it says “done” and the test still fails → loop and control flow
- Break 4: the cheapest way to pass a test is to delete it → permissions
- Break 5: you fixed four things and can’t tell if it’s better → evals
- How the pieces interact
- Where this framing has limits
- Check yourself
- Famous related terms
- Going deeper
The picture version
The whole idea in six pictures, for a reader who has never built an agent. The prose below fills in the seams the pictures skip.
1 · The problem
Same brains. Wildly different products.
2 · The naive way
Build the worst possible harness first: ten lines.
3 · The walk
Five things break, in order. Each break names a layer.
4 · The catch
Score it on a green suite and deleting the test wins.
5 · The measure
One green run tells you almost nothing.
6 · Keep this card
The whole thing on one index card.
Why it exists
You ask a coding agent to fix one failing test in your repo. It reads the test, opens the file it points at, makes an edit, re-runs the suite, sees a different failure, tries again — and eventually hands you a green build. Then you paste that same failing test into a plain chat window pointed at the same model family, and you get a confident code block you have to apply yourself, against a file the model never actually read.
You probably assume the gap is the model — that the agent is running something smarter. It usually isn’t. What differs is everything around the model: which tools it can call, which files end up in its prompt, how the loop decides when to stop, what gets retried after a failure. That wrapping is the harness. A model alone is like a brain in a jar — smart, but with no eyes, no hands, no memory of yesterday. The analogy breaks in one important place: the brain in the jar at least stays the same brain. A model is re-rolled every call, reconstructing its entire sense of the situation from whatever text the harness hands it. “Harness engineering” is the (newly named) craft of building bodies for a worker with no memory of its own.
That failing test is the example this post keeps coming back to.
Notice what this implies about the current landscape. The same handful of frontier models — a few families, all trained by labs you can name — sit inside a great many very different-feeling products. One feels like a tireless pair programmer. One feels like a research assistant that browses the web patiently for an hour. One feels like a scatterbrained chatbot that forgets what you said three turns ago. Same brains, wildly different behavior.
The thing that varies isn’t the model. It’s the harness — the program that runs around the model, deciding what tools it has, what stays in its context, when to stop, what to retry, what to ask the user before doing. Harness engineering is the discipline of designing, measuring, and improving that program.
It exists as a named thing because the work turned out to be its own craft. You can’t do it well by being a good ML engineer; the model is a black box you’re prompting, not a thing you’re training. You can’t do it well by being a good systems engineer either; your “system” is non-deterministic and re-rolls the dice every run. The skill set is something else: part product design, part distributed-systems-with-an-unreliable-worker, part prompt craft, part eval design. Hence its own name.
Why it matters now
Almost every AI product you actually touch is a harness: the coding agent that fixes your failing test, the customer-support bot that has to look up your order before answering, the “deep research” tool that browses for an hour, the computer-use agent clicking through a form. In each, the model is bought off a shelf a handful of labs stock; the product is the wrapping. My read — a working thesis, not a measured claim — is that model choice is no longer the dominant variable in how good most of these feel. It shows up in a few shapes worth taking seriously:
- The same model in two different harnesses can behave like two different products. A coding agent with carefully designed tools, context pruning, and a verifier loop will often out-ship a coding agent on a stronger model with a sloppy harness. Common folklore among teams who ship agents; not a measured result.
- Many user complaints are about harness, not model. “It forgot what I told it” — context management. “It deleted the wrong file” — permissions. “It looped forever” — stopping condition. “It hallucinated an API” — tool design and verification. The model is often not the proximate cause; it’s the thing the harness failed to corral.
- Reliability gains live here. The fastest way to make an unreliable agent more reliable is rarely “wait for the next model.” It’s tighter tools, shorter horizons, better verifiers, cleaner context. Your failing-test agent gets more reliable when the harness runs the suite itself instead of believing the model’s “fixed it.” (See Why agents fall apart over long horizons for the structural reason.)
This is also why “prompt engineering” stopped being the whole story. Prompt engineering is one slice of harness engineering — the system-prompt slice. The rest of the harness — tools, loop shape, memory, permissions, eval — matters at least as much, often more.
The short answer
harness engineering = tool design + context engineering + loop & control flow + permissions + evals, all iterated against a non-deterministic worker
Picture to keep: a brilliant contractor with total amnesia who shows up every morning knowing nothing — so the whole job depends on the briefing folder you hand them, the tools you leave on the bench, the gate you put in front of the demolition equipment, and someone checking the work before the client sees it. Harness engineering is designing that workplace. Where the analogy breaks: a real contractor gets better at your house over months. This one never does — every improvement has to be built into the workplace, because none of it accumulates in the worker.
It’s the craft of building everything around a language model so that the combined system is useful, safe, and improvable. The model is the worker; the harness is the workplace, the manager, the safety officer, and the QA team.
How it works
The fastest way to understand a harness is to build the worst possible one and watch it fail. Start with the whole thing in ten lines:
while not done:
reply = model(conversation)
conversation += reply
conversation += run_any_tools(reply)
Point that at “fix the failing test.” Every layer of real harness engineering is the fix for something that breaks in the next few minutes.
Break 1: it has no hands → tool design
The loop above can only talk. Give it read_file, write_file, and
run_tests and it can actually work the problem. But now the model’s whole
interface to your repo is a set of function signatures — and it turns out
those signatures behave like part of the prompt on every subsequent step.
Their names, descriptions, argument schemas, and — crucially — their
error messages all condition what it does next.
The unintuitive part: a model with great tools behaves like a smarter model. A model with bad tools behaves like a dumber one. Concretely:
- Names and descriptions matter as much as code.
read_filevs.fetch_path_contentsis not a stylistic choice — in practice the model often picks tools partly by lexical match against the task wording. Inconsistent vocabulary across tools tends to cost accuracy. (Operator heuristic, not a measured effect.) - Error messages are teaching signals. If
edit_filefails, the string it returns is what the model will read and condition on. “File not found” is fine; “File not found. Did you mean to create it? Usewrite_filefor new files.” is better — it routes the model’s next action. - Schema strictness is a design choice, not a default. Loose schemas let the model pass the wrong types; over-strict schemas burn turns on validation errors. The right level depends on how forgiving downstream code is.
- Tool surface area has a cost. Every tool you add competes for attention in the system prompt and broadens the space of things the model might wrongly try. Ten well-chosen tools usually beat fifty exhaustive ones.
Break 2: thirty steps in, it forgets the job → context engineering
Good tools, and the loop runs. By step thirty the conversation holds four file dumps, six test outputs, and three abandoned attempts — and the agent starts re-trying a fix it already watched fail, or “fixes” a file that has nothing to do with the original test. Nothing about the model changed; the thing it reads changed.
That’s the constraint the naive loop ignored. Models have a finite context window, and even within that window, attention isn’t uniform — material in the middle of long contexts is often used worse than material at the edges (the lost-in-the-middle effect). Because the loop blindly appends, the original goal ends up buried in the worst possible position. So “what’s in the context, in what order, in what shape” is a real design problem.
The standard moves:
- Summarize old turns once they’re far enough back that detail no longer matters. The harness — not the model — decides when.
- Retrieve on demand instead of pre-stuffing. If the agent can call
read_file, you don’t need to dump the whole repo into the prompt. - Pin invariants at the top: the user’s actual goal, the constraints that must hold, the plan. These get re-read every step.
- Prune contaminated history when the agent has gone down a wrong path. Leaving every failed attempt in context is exactly the substrate that self-conditioning feeds on.
- Cache aggressively. If the system prompt and tool definitions don’t change between calls, prompt caching can shift a large fraction of each turn’s cost into a one-time charge — exact savings depend on the provider, the cache TTL, and how much of the prefix is stable. (See Why prompt caching exists.)
The hard part is that these moves trade off. Summarizing too eagerly loses the detail the next step needs; pruning too aggressively erases the reason a path was rejected and the agent re-tries it.
Break 3: it says “done” and the test still fails → loop and control flow
Look again at the naive loop’s first line: while not done. Who decides
done? In the ten-line version, the model does — it stops emitting tool
calls and declares victory. Sometimes the test is green. Sometimes it
edited the assertion instead of the code. Sometimes it never stops at all,
re-running the suite forever.
So the loop is where the harness stops trusting the model’s self-report. The interesting questions are everything that goes around those ten lines:
- When does the loop end? Model says “done”? A budget hit? A verifier passes? Realistic harnesses have several stopping conditions and pick the first one to trip.
- Plan first or improvise? Plan-then-execute bounds how far one bad step propagates, at the cost of flexibility. Pure ReAct-style improvisation is more flexible but compounds errors faster.
- One agent or several? Sometimes the right move is one big loop; sometimes it’s a coordinator that spawns a fresh sub-agent per subtask, each with its own clean context. Sub-agents are the closest thing harness engineering has to a “fork the process” primitive.
- What runs after the model speaks but before the user sees it? Linters, type-checkers, test runners, schema validators, policy checks. These are cheap and brutal — the agent thinks it’s done; the harness disagrees, and the loop continues.
Break 4: the cheapest way to pass a test is to delete it → permissions
Add a verifier and you’ve told the agent exactly what it’s being scored
on. A model that can call write_file on anything now has a very cheap
path to a green suite: delete the test. This is not hypothetical
malice — it’s the loop optimizing the target you gave it, using a tool you
handed it, in a directory you didn’t fence off.
The fix isn’t a better prompt. The model is allowed to propose anything; the harness decides what actually executes. Concretely:
- Read vs. write asymmetry. Most harnesses let the model freely read files, search, and inspect, but pause before writes, deletes, network calls to production, or anything that costs money.
- Allowlists over denylists. “These tools can run without asking” is more robust than “these tools require confirmation.” With a denylist, any tool you didn’t think to flag as dangerous defaults to running silently — that’s where unknown unknowns live. With an allowlist, the default for anything new is “ask,” which is the safer failure mode.
- Confirmation UX is part of the harness. A confirmation prompt that buries the relevant detail (which file? what diff?) gets rubber-stamped and provides no real safety. A clear one is the difference between a rail and a placebo.
This is also where the harness’s relationship to the user lives. “Auto mode” vs. “ask before each step” isn’t a UI toggle layered on top — it’s a fundamental knob in the harness itself.
Break 5: you fixed four things and can’t tell if it’s better → evals
Now the real problem. You’ve changed tool descriptions, added pruning, put in a verifier, locked down writes. You re-run “fix the failing test” and it works. Did any of it help? Run it again and it might fail — the agent is stochastic, and the run you just watched is a sample of one. Worse, the context pruning you added may have quietly broken a different task by deleting the reason a path was rejected.
This is where harness engineering looks least like model training and most like its own thing. One good run and one bad run on the same task tell you very little on their own, so the honest workflow is closer to A/B testing than to unit tests:
- A frozen task suite, ideally tasks that take the agent more than a trivial number of steps so harness effects show up.
- Multiple runs per task per variant — a single sample is noise.
- Metrics that distinguish “got the right answer” from “got there cleanly” — token cost, tool-call count, wall-clock time, number of retries, human interventions per task.
- Traces, not just logs. When something goes wrong on step 27 of 40, the only debugging substrate is the full transcript: model inputs, outputs, tool results, decisions the harness made. A trace-viewing UI tends to get built or bought early, for the simple reason that the alternative is reading step 27 of 40 as raw JSON. (How common that is isn’t something public data settles — it’s an impression from open tooling and write-ups.)
A specific failure mode worth naming: eval drift on real tasks. The production workload changes faster than your eval suite, and your evals slowly stop reflecting it. The discipline is rotating in fresh tasks from real user traces — anonymized — at a steady cadence.
How the pieces interact
The reason these aren’t independent: a change in one layer often only pays off if another layer changes too.
- Adding a verifier (loop layer) is wasted if the harness can’t act on its signal — i.e., can’t roll back a contaminated state (context layer).
- A tool with great error messages (tool layer) is wasted if those errors get summarized away three turns later (context layer).
- Tighter permissions (safety layer) without clear confirmation UX (loop / UX layer) just train users to click through.
The mental model that works: harness engineering is iteration on a non-deterministic compound system, where every change has to be evaluated end-to-end because local improvements can degrade global behavior in surprising ways.
Where this framing has limits
A few honest caveats:
- “Harness engineering” is an emerging label. People have been doing this work since the first tool-using agents; the name is recent, and it spread through practice, source code, and blog posts rather than a paper. If you read this in a few years, the boundaries between “harness engineering,” “AI engineering,” and “agent design” may have settled differently.
- The model still matters. A weak model in a perfect harness is still a weak product. The claim is that, between today’s frontier models, the harness is the dominant variable — not that models are irrelevant.
- Some tasks are model-bound. A task that fails because the model genuinely can’t reason about the domain isn’t going to be saved by better tool design. Knowing which kind of failure you’re looking at is itself a harness-engineering skill.
- The discipline is young and the literature is thin. Most of what’s known about harness engineering is folklore inside teams that have shipped agents, blog posts, and source code of open agents. There isn’t a textbook yet, and a lot of strong-sounding claims (including some in this post) are working hypotheses, not measured results.
You started this post with a ten-line loop and a failing test. What did the
five breaks add? — + tools + context + a stopping rule + a fence + a way to measure, and the last one is the one that makes it engineering rather
than tinkering. Without evals you can still change all four other layers;
you just can’t tell whether you improved the agent or got a lucky run.
Check yourself
Before you go — an agent keeps “fixing” your failing test by editing the
test file. You add a permission rule blocking writes to tests/. It now
gets stuck, burning fifty steps and giving up. Which layer did you actually
break?
Answer
The loop layer, via the context layer. Blocking the write was right, but
now the agent’s context fills with rejected-permission errors it can’t act
on, and nothing tells it “that path is closed, solve the real bug instead.”
A rail without a route just converts a wrong answer into an expensive
non-answer. The fix lives in the tool’s error message (“writes to tests/
are blocked; the assertion is correct — fix the implementation”) and in a
step budget that stops the bleeding. This is the interaction effect from
the section above: a safety change that only pays off if the tool and loop
layers change with it.
And one more — you swap your agent onto a model that scores meaningfully higher on coding benchmarks, and end-to-end task success barely moves. What would you look at before concluding the new model isn’t better?
Answer
Whether the failures are model-bound or harness-bound. Read traces, not scores: if the runs die on timeouts, permission stalls, context exhaustion, or tools called with the wrong arguments, a smarter model has no room to express itself — you’re measuring your harness’s ceiling, not the model’s. The diagnostic question is “at the step where this went wrong, did the model reason badly, or did it reason fine about bad inputs?” Only the first kind of failure gets fixed by a better model.
Famous related terms
- Agent harness —
agent = model + harness— the noun this discipline is the verb of. - Prompt engineering —
prompt engineering = one slice of harness engineering— historically the whole story; today one sub-skill alongside tool design, context engineering, loop design, and eval. - Context engineering —
context engineering = deciding what's in the model's context, in what shape, at what cost— the sub-discipline most teams discover second. - Tool / function calling —
tool use = model emits structured calls + harness executes them— the protocol via which the harness offers the model things it can do. - MCP (Model Context Protocol) —
MCP = an open protocol for exposing tools, resources, and prompts to AI applications— lets you reuse harness work across products. - Eval harness —
eval harness = task suite + runner + scoring— overloaded term: in research, it’s the rig that scores models on a benchmark; in product work, the rig that scores your harness on your task suite. Same word, different artifact. - Scaffolding —
scaffolding ≈ harness— an older word for roughly the same idea, and still used interchangeably; my impression is that it reads as the more academic of the two.
Going deeper
- Anthropic, Building effective agents (December 2024) — first-party guidance for “which loop shape should I build?”, naming the patterns (chaining, routing, orchestrator-workers) this post’s control-flow section gestures at, with the caveat Anthropic itself now adds: the loop shapes have aged better than the surrounding tooling advice.
- The source of an open coding agent — Aider is a readable one — is the explainer no article can be: it answers what a harness literally looks like as code, and how much of it turns out to be prompt strings rather than control flow.
- Sinha et al., The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs (2025) — the rabbit hole, for “why does step 40 go so much worse than step 4?”; it studies long-horizon execution and self-conditioning rather than harness design, but it’s what makes verifiers and context pruning feel necessary rather than fussy.