edidiong umana · writing
home
Engineering notes6 min read

The harness is the product: an agent is the model plus everything around it

Before you pay for a bigger model, fix what surrounds it. A practical tour of context files, skills, tools, permissions, sandboxes and checks, with a minimal AGENTS.md you can copy into your repository today.

Your coding agent says "Done, all tests pass." You open the app and the main feature doesn't work. The obvious move is to switch to a bigger, more expensive model. Often the cheaper fix is everything around the model.

Anthropic's engineers ran a clean comparison. They asked a model to build a retro game maker from a one-sentence prompt, twice. Running solo, it took 20 minutes and $9, and the central feature, playing the game you had built, was broken. In a full harness, a planner wrote a spec, a builder worked through it sprint by sprint, and a separate evaluator clicked through the app and sent bugs back. That took 6 hours and $200, and the core worked. Same model. The environment decided the outcome, and it was a trade: reliability bought with time and money.

Agent = model + harness

LangChain's engineers put the definition simply: an agent is a model plus a harness, and the harness is all the code, configuration and logic that isn't the model. That means the instruction files, the tools and their descriptions, the sandbox they run in, how long histories get summarised, what the agent may do without asking, and the checks that run before an answer counts.

The evidence that this half matters is real. LangChain lifted its coding agent on Terminal-Bench 2.0 from 52.8% to 66.5% with the model fixed, changing only the harness. But read it carefully. In the same harness, two different models scored 11.5% and 82.2%, so the model still matters most. And harness changes can hurt: in April 2026 Anthropic traced a drop in Claude Code's quality to three harness changes, one of them a single line in the system prompt.

No primary source measures a split like "the harness is 90% of performance". The useful habit is narrower: when an agent fails, change its environment instead of asking it to try harder, and measure one change at a time on your own tasks.

AGENTS.md: a map, not a manual

AGENTS.md is an open format: a plain Markdown file at the root of your repository that coding agents read at the start of every session. Codex, Gemini CLI, OpenCode, goose, Aider, Cursor and GitHub Copilot support it, and since December 2025 it sits with the Linux Foundation's Agentic AI Foundation. Claude Code reads its own CLAUDE.md and can read AGENTS.md too.

Most teams make it long. OpenAI's Codex team tried that first and found it crowded out the task, went stale, and couldn't be checked. Their fix was a file of about 100 lines that works as a table of contents, pointing into a docs/ folder that holds the real knowledge.

An ETH Zurich preprint adds a warning. Context files raised agents' costs by 20 to 23%. Files written by an LLM slightly lowered task success, and files written by developers raised it by 2.4%, which was not statistically significant. Repository overviews didn't help. So write it yourself, and only for what the agent can't discover: exact commands, rules the code doesn't make obvious, and pointers. Leave style to the linter. Here is a minimal one:

# AGENTS.md

Bill-splitting API for savings groups. Python 3.12, FastAPI, SQLite.

## Commands
- Set up: `./init.sh` (creates .venv and installs dependencies)
- Check: `python check.py` (about 10 seconds). Run it before you say you are done.
- Full tests: `pytest -q` (slow; CI runs them on every push)

## Easy to get wrong
- Money is a whole number of kobo. Never use floats.
- Never edit files in `migrations/`. Create a new migration instead.
- Never edit a test to make it pass. Fix the code.

## Where to look
- `docs/architecture.md`: how a request flows, and which module may import which
- `docs/payments.md`: read before touching anything in `app/payments/`
- `progress.md`: read it and `git log -5` when you start; add two lines when you finish

Mitchell Hashimoto's rule keeps it honest: add a line only after you have watched an agent make the mistake it prevents.

Skills and just-in-time context

Your agent doesn't need the deployment runbook on every turn, only on the day it deploys. A skill is a folder with a SKILL.md file: a short header with a name and a description, then instructions, plus any scripts the procedure needs. Anthropic published Agent Skills as an open standard in December 2025, and Claude Code, Codex and other harnesses read the same format.

Skills load in stages. At startup the harness reads only each name and description. The full instructions load when a task matches, and scripts only when the instructions call for them. A project can carry dozens of skills and pay about a line each until one is used. So write the description as when to use the skill, not just what it is.

The same thinking applies to tool output, which is often the biggest part of the context. Anthropic's advice is fewer, larger tools, concise results by default, and truncated or paginated big ones: in one of its examples, a concise response took 72 tokens and a detailed one 206.

Tools, permissions and sandboxes

An agent that can run commands can also delete files or leak keys, and anything it reads can carry instructions: a web page, an issue, a README. Simon Willison calls the dangerous combination the lethal trifecta: private data, untrusted content and a way to send data out. His advice is to avoid the combination, because filters can't be trusted to stop it.

Anthropic's conclusion is to contain first, then steer. In one test, setup steps pasted in by a user hid a request to send cloud credentials to an outside server, and the agent did it in 24 of 25 tries. If credentials never enter the agent's environment, they can't leak. In practice:

  • Permission rules. Allow the safe, frequent commands such as your check script. Deny what must never happen, such as reading .env or pushing to git. Everything else asks.
  • A sandbox. Operating-system limits on files and network, enforced whatever the model decides. Inside Anthropic, Claude Code's sandbox cut permission prompts by 84%. Codex CLI runs with network access off by default.
  • Test keys only, with spending limits. Never production keys.

Don't lean on approval prompts: Anthropic found people approve 93% of them. And a line in an instruction file is not a control. A study of 481 CLAUDE.md files found only about 4 to 16% of their security rules backed by a real setting.

Feedback loops the agent can't argue with

Birgitta Böckeler, writing on Martin Fowler's site, sorts the harness you build into guides, which steer before the agent acts (AGENTS.md, skills, docs), and sensors, which check after it acts (tests, type checkers, linters, review agents). Prefer sensors that are plain code: fast, cheap and the same every time.

Agents are confident, and confidence isn't evidence. Anthropic found that agents grading their own work tend to praise it. A 2026 preprint on agent loops found an agent claiming an improvement in all 54 cycles, while in 56% of them nothing had really improved. So put the check outside the agent's opinion:

  • The harness runs the checks before it accepts "done". In Claude Code, a stop hook that exits with code 2 sends the agent back to work.
  • Quiet when green, loud when red. HumanLayer estimates that 200 lines of passing test output can use 2 to 3% of a context window for nothing. Print one line on success.
  • Errors that teach the fix. OpenAI's custom lint messages tell the agent how to fix the problem, not just what broke.

Most of what LangChain changed for its 13.7-point gain was this kind of back-pressure: a verification step, a checklist before finishing, loop detection and time-budget warnings. None of it made the model smarter. It made it harder to stop early.

Do this todayTake a task your agent recently got wrong. Add the exact test command to your AGENTS.md with the line "Run it before you say you are done", then run the same task in a fresh session and compare.

Learn it properly

This post is a slice of Track 5 of the AI Study Group, Harness engineering. It's free, works with local models, and ends with a lab where you build a small harness in Python that refuses to accept "done" until the tests pass. The track follows the arc of WalkingLabs' open-source course Learn Harness Engineering (MIT licence), rewritten with every number re-checked against its primary source.

Sources