aliteq.

AI agent keeps making the same mistake? Harness engineering, explained

The model is only one part of an agent. The loop, the tools, the permissions and the memory around it are the harness, and in 2026 OpenAI, Anthropic and Mitchell Hashimoto all described fixing agents there instead of in the prompt. Here is what that means and where the name came from.

KernelUpdated 11h ago11 min readWeb story
Hand-drawn editorial illustration of a small coral robot walking along a path inside an off-white strap harness held by human hands, on a deep near-black background with indigo-violet shapes
Share

Think of a workshop. The model is a very fast apprentice. The harness is the bench, the labeled tool wall, the lockable cabinet and the checklist taped above the door. Same apprentice, very different results depending on the shop.

I am not a person who has built one of these at scale, so this page does what I do best: it reads the primary sources and puts them side by side. If you want the beginner version of related ideas, start with context engineering for vibe coders and what a context window is.

What is harness engineering?

Harness engineering is the work of designing everything around a model that makes it behave as an agent: the loop, the tools, the rules about what it may touch, the context it sees and the tests that grade it. Anthropic defines an agent harness as the system that processes inputs, orchestrates tool calls and returns results.

That definition comes from Anthropic's evals post, read 3 October 2026. A model alone only returns text. A harness gives it hands (tools), a short-term memory (context), a boss (permissions) and a report card (evals). Take any one away and you get a chatbot, a runaway script, or something you cannot trust.

Two things it is not. It is not prompt engineering, which tunes the words in one request. And it is not the model itself. When a new model arrives, a good harness keeps most of its value because the scaffolding is yours.

Who coined "harness engineering"?

Nobody can prove it from public sources. Hashimoto titled a step "Engineer the Harness" in an essay dated 5 February 2026. OpenAI published "Harness engineering: leveraging Codex in an agent-first world" on 11 February 2026. Anthropic was already calling the Agent SDK an "agent harness" in November 2025.

Here is the timeline I could confirm on the primary pages:

  • 26 Nov 2025: Anthropic's long-running-agents post describes the Claude Agent SDK as "a powerful, general-purpose agent harness adept at coding."
  • 9 Jan 2026: Anthropic's evals post separates the agent harness from the evaluation harness.
  • 5 Feb 2026: Hashimoto writes that anytime an agent makes a mistake, you take the time to engineer a solution so it never makes that mistake again.
  • 11 Feb 2026: OpenAI, by Ryan Lopopolo, puts "Harness engineering" in a post title.

Several blogs say OpenAI coined the term. I could not verify that on any primary source, so I will not repeat it. What I can say: the word "harness" was already in use, Hashimoto named the habit first by date, and OpenAI's post is the one that made it a discipline with a name.

What does Hashimoto actually mean?

His version is a habit, not a framework. Every time an agent makes a mistake, fix the environment so that mistake cannot recur. He describes two ways: better written instructions in a file such as AGENTS.md, and small programmed tools that let the agent check its own work, such as screenshot scripts and filtered test runners.

On his Ghostty project, he writes that each line in the instructions file "is based on a bad agent behavior." That is a useful test for your own rules file. If a line does not trace back to a real failure, it is probably noise.

What does the loop look like?

An agent loop has four beats. The harness hands the model the prompt, the tool list and the history. The model answers or asks for a tool. The harness runs the tool and returns the result. That repeats until the model replies with no tool calls. Each full cycle is one turn.

Four steps of one agent-loop turn: receive the prompt, tool list and history; the model evaluates and answers or requests a tool; the harness executes the tool subject to hooks and permission rules; results feed the next turn until a reply has no tool calls. Example: fixing failing tests took four turns, three with tools and one text-only.
One turn of the loop, from the Claude Agent SDK docs, read 3 Oct 2026. · aliteq research

Anthropic's docs walk through "Fix the failing tests in auth.ts": run the tests, read two files, edit and re-run, then a text-only answer. That is four turns. A bigger task can chain dozens.

Anthropic's Agent SDK post summarizes the working pattern as gather context, take action, verify the work, repeat. The verify step is the one people skip, and it is where harness work pays off most.

Two controls matter for cost. max_turns counts tool-use turns and max_budget_usd caps spend, and the docs say both default to no limit. The docs add that setting a budget is a good default for production agents.

What tools should the harness give an agent?

Give the agent a small set of tools that do real work and say clearly what each one does. Anthropic's post calls tools "the primary building blocks of execution." The Agent SDK ships file tools, search tools, a shell tool, web tools and tools for spawning subagents, and you can add your own or connect MCP servers.

Two harness details are easy to miss. Read-only tools can run at the same time, while tools that change state (edit, write, shell) run one after another to avoid conflicts. And every tool definition costs context, so a few servers with many tools can use up a lot of context before the agent does anything. The docs say tool search now defers MCP tool schemas by default.

OpenAI's team went further than tools for files. They made the app bootable in each git worktree, wired in browser tooling and exposed logs and metrics so the agent could query them. Their stated goal was making the app, logs and metrics legible to the agent.

How do permissions work in a harness?

Permissions are rules about what the agent may run without asking. In Anthropic's Agent SDK you combine three things: an allow list that auto-approves tools, a deny list that blocks tools regardless of other settings, and a permission mode that sets overall oversight. A denied call goes back to the model as a result, and it usually tries another way.

Scorecard of six Claude Agent SDK permission modes: default asks a callback; acceptEdits auto-approves edits; plan never edits source; dontAsk denies anything not pre-approved; auto uses a model classifier; bypassPermissions runs everything allowed and is for isolated environments only.
The six permission modes in the Agent SDK docs, read 3 Oct 2026. · aliteq research

You can also scope a rule, such as allowing only shell commands that start with npm. And hooks run code before or after a tool call. A hook that rejects a call stops it, and the docs note hooks run in your app, not in the model's context, so they cost no tokens.

The docs say to use the bypass mode only in isolated environments where the agent's actions cannot affect systems you care about. If you ship apps built with an AI tool, my colleague's vibe-coded app security checklist covers the human side of the same problem.

How should a harness manage context?

Treat context as a scarce budget. Anthropic describes context engineering as curating the smallest useful set of tokens at each step. It warns about context rot, where recall gets worse as the window fills. The usual fixes are summarizing old history, writing notes outside the window, giving subtasks to fresh subagents and loading data only when needed.

In the Agent SDK, compaction summarizes older history near the limit, and the docs warn that early instructions "may not be preserved," so durable rules belong in a project file like CLAUDE.md. A subagent starts clean and sends back only its final answer, so the main context grows by a summary instead of a transcript.

Anthropic's long-running-agent post shows the same idea across sessions. Each new session "begins with no memory of what came before," so the harness leaves a trail: an init script, a progress file and a feature list the next session reads first.

OpenAI's lesson was blunt. They tried one big AGENTS.md and it failed: it crowded out the task, it went stale and it was hard to verify. They moved to a roughly 100-line AGENTS.md that works as a table of contents, pointing into a structured docs folder. Their phrase: "give Codex a map, not a 1,000-page instruction manual." Another line is worth pinning up: anything the agent cannot access in context while it runs effectively does not exist.

What do evals have to do with the harness?

Evals tell you whether a harness change helped. Anthropic's vocabulary: a task is one test with success criteria, a trial is one attempt (run several, since models vary), and a grader scores the result. Graders can be code, another model or a human, and each trades speed against nuance.

Code graders are fast and repeatable but can be brittle. Model graders are flexible but non-deterministic and need calibration. Humans are the gold standard but slow and costly. Anthropic also separates two measures: pass@k is the chance of at least one correct answer in k tries, and pass^k is the chance that all k succeed. For anything customer-facing, the second one is the honest number.

Note the two meanings of "harness." The agent harness runs the agent. The evaluation harness runs the tests that grade it. They are separate pieces of infrastructure.

What did OpenAI's experiment show?

OpenAI reports that over five months a small team shipped an internal beta with zero manually written code, about a million lines and roughly 1,500 merged pull requests. The team says it spent its effort on environment, constraints and feedback loops, not typing. It also reports a cost: cleaning up "AI slop" took 20% of each week until they automated it.

OpenAI's reported numbers for its agent-built repository: zero manually written lines, about 1,500 merged pull requests, 3.5 pull requests per engineer per day, an AGENTS.md of roughly 100 lines, and 20 percent of each week spent on cleanup before automation.
Figures as OpenAI reported them on 11 Feb 2026, read 3 Oct 2026. Not independently checked. · aliteq research

What they built is the useful part, because it is a catalog of harness pieces:

  • Strict architecture, enforced by machines. Fixed layers with checked dependency directions, enforced by custom linters and structural tests. OpenAI wrote the lint error messages so they inject fix instructions into the agent's context.
  • A docs folder as the source of truth, kept honest by CI checks and a recurring "doc-gardening" agent.
  • Garbage collection. Because the agent copies patterns it finds, including bad ones, they encoded "golden principles" and ran recurring background cleanup tasks that open small refactor pull requests.
  • Light merge gates, a choice OpenAI says would be irresponsible in a low-throughput environment.

Two caveats from OpenAI itself. It says the full autonomy "depends heavily on the specific structure and tooling of this repository and should not be assumed to generalize without similar investment." And it says it does not yet know how architectural coherence holds up over years. Read the headline numbers with that in mind.

How do you start with harness engineering?

Start with one real failure. Write down what the agent got wrong, then ask what was missing: a rule, a tool, a check or a piece of context. Add that, in code if you can, and run the task again. This is Hashimoto's habit and OpenAI's loop in miniature.

Cap the loop. Set a turn limit and a spend limit before you let an agent run unattended.

Start strict. Allow only the read-only tools it needs, deny the rest, and widen one tool at a time.

Write a short rules file. Each line should trace to a real mistake. Keep it a map that points elsewhere, not an encyclopedia.

Give it a way to check itself: a test command, a linter, a screenshot. Then add a small eval with several trials per task.

Turn repeated review comments into lint rules or hooks, so the rule runs every time and does not depend on the model remembering it.

If you build with an AI tool rather than writing agents, the same idea scales down: a good rules file and a test command are your harness. The shift from typing code to steering agents is also the subject of our pieces on Karpathy's agentic engineering and vibe coding versus agentic engineering. For more agent coverage, see the AI agents hub and our guide to Claude Code for vibe coders.

Quick answers

What is harness engineering in simple terms?
It is building everything around an AI model so an agent works reliably: the loop that calls it, its tools, permission rules, context handling and evals. When the agent makes a mistake, you fix the surrounding system so it cannot repeat it. OpenAI's 11 Feb 2026 post helped popularize the name.
Who coined the term harness engineering?
I could not verify one originator. Mitchell Hashimoto wrote "Engineer the Harness" on 5 Feb 2026, OpenAI titled a post "Harness engineering" on 11 Feb 2026, and Anthropic used "agent harness" in Nov 2025. Claims that one company coined it were not on any primary source I read.
Is a harness the same as a framework or an SDK?
Not exactly. An SDK such as Anthropic's Agent SDK ships a general-purpose harness you can configure. Your project's rules file, tool scripts, permission settings, hooks and evals are the part you engineer on top. Anthropic calls the SDK itself "a powerful, general-purpose agent harness."
How is harness engineering different from prompt engineering?
Prompt engineering tunes the words of one request. Harness engineering changes the system around the model: which tools exist, what is allowed, what context loads, what gets checked. OpenAI described fixes as capabilities that are "legible and enforceable" for the agent, not as "try harder" prompts.
Do I need to be a developer to use these ideas?
No. If you use an AI coding tool, a short rules file, a test command and sensible permission settings are a small harness. The agent-building details here, like turn limits and permission modes, matter if you write agents with an SDK.
Does a better harness mean I can trust the agent without review?
No. OpenAI says its autonomy depended on repository-specific tooling and should not be assumed to generalize. Anthropic's docs tell you to keep the loosest permission mode for isolated environments. A harness lowers risk; it does not remove it.

Found this useful? Share it

Share
Kernel

Software & Business Software Editor

Kernel

I'm US-based, I've daily-driven more Linux distros than I can name, and I treat software like a workshop: what does it do, what does it really cost, and what can I run myself instead. That's why I also cover the CRM, HR and ERP bills that land on a startup the day it signs its first big customer.

The Aliteq brief

The tech worth knowing — hardware, AI, gaming, deals. No spam, unsubscribe anytime.

Keep reading