Harness Engineering: Making Coding Agents Dependable

Instruction files that actually bind


Ledgerly's first instruction file was 310 lines long. Someone pasted in the team wiki: the architecture overview, GST rules, the deploy process, a style guide and, on line 214, the command for running the fast tests. Agents still ran the 41-second full suite after every edit. The one line that would have saved half an hour per task was there. It was buried under 4,000 tokens of text that did not matter for the task at hand.

Every major coding agent reads an instruction file at the start of a session. Claude Code reads CLAUDE.md. Codex, Cursor and a growing list of other tools read AGENTS.md. The mechanism is the same everywhere: the file's text is placed into the context before your task. That makes it powerful, because the agent sees it in every session. It also makes it expensive, because every line competes for the model's attention with the task itself.

This lesson is about writing a file that the agent actually follows, and about knowing where an instruction file's power ends.

Three strengths for one ruleInstruction file — says what and whyRoute — deeper docs, read when neededGate — what must be true at the endPermission — what must never happen
The migration rule in the instruction file cut bad edits from 6 in 20 runs to 2; only the permission rule made it zero.

Root, route, check

A good instruction file does three jobs, and only three.

  • Root: the few facts that every session needs. What the project is, where things live, the commands, and the handful of rules whose violation causes damage.
  • Route: pointers to deeper documents that only some sessions need. "Before changing tax code, read docs/gst.md." The agent reads that file only when the task goes there.
  • Check: every rule should be something the agent or the harness can verify. "Write clean code" cannot be checked. "ruff check ledgerly tests passes" can.

Here is Ledgerly's rewritten root file. It is 26 lines, about 450 tokens.

Text
# Ledgerly: instructions for coding agentsLedgerly is a Flask app for GST invoices and payment reminders (Python 3.11).Code is in ledgerly/, tests in tests/, database migrations in migrations/.## Commands- Fast tests while you work (9 s):     python -m pytest tests/unit -q- A single test file:                  python -m pytest tests/unit/test_fees.py -q- Full check before you finish (45 s): bash scripts/check.sh- Lint only:                           ruff check ledgerly tests## Rules- Money is Decimal, never float. Amounts are rounded to 2 places, half-up.- Every behaviour change needs a test in tests/unit that fails without the change.- Never edit a migration that is already merged. Create a new one instead.- Do not open or print .env. Tests get settings from tests/fixtures/settings.py.- tests/integration/test_pdf_export.py is flaky (about 1 run in 8). It is marked  flaky and excluded from check.sh. Do not skip, delete or "fix" it.- Work only on your task. Report unrelated problems in your final message,  one per line, starting with QUEUE:.## Where to look- Fees and rounding: ledgerly/fees.py. Tax: ledgerly/tax.py, and read docs/gst.md first.- Invoice statuses and transitions: docs/invoice-states.md.- ledgerly/utils.py is legacy. Read only the function you need. Do not refactor it.- Migrations: read migrations/README.md before creating one.

Look at what is not in it. No architecture essay: the agent can read the code. No list of libraries: pyproject.toml has that. No deploy process: the agent never deploys. No GST rules: they sit behind a route and cost nothing on tasks that do not touch tax.

Keep, move or delete

To shrink an old file, go through it line by line and give every line one of three verdicts. The test for each line is: if this line were missing, would a session go wrong in a way I can name?

Line from the old 310-line fileVerdictWhy
"Write clean, maintainable code."DeleteNot checkable, and the model already tries
"We use Flask 3 and SQLAlchemy 2."DeleteVisible in pyproject.toml
"Run fast tests with python -m pytest tests/unit -q."Keep, move to the topSaves about 30 minutes per task
"Never edit merged migrations."Keep, and enforce in the harnessToo costly to leave to a request
40 lines explaining GST rulesMove to docs/gst.mdNeeded only by tax tasks
25 lines on the deploy processDeleteAgents never deploy Ledgerly
"Be careful with utils.py."RewriteSay what to do: read one function, no refactors
"IMPORTANT: ALWAYS run ALL tests after EVERY change!!!"DeleteContradicts the fast-test rule; capitals do not help

The last row is common and worth a word. Shouting in capital letters is a sign that an earlier rule was ignored. The usual reason it was ignored is that it conflicted with another rule or was buried. Fix the conflict or the position. Louder text just adds noise, and current models tend to over-apply emphasised rules to cases where they do not fit.

Two other signs that a line should go: it describes something that is already true of the code (the model will see it anyway), or nobody can remember which failure it was added to prevent.

Where instruction files live

Most tools support more than one instruction file. A root file loads at the start. Files in subdirectories load when the agent works in that folder, and for AGENTS.md the nearest file generally takes precedence. This is routing done by the tool for you.

Use it for rules that only matter in one place. Ledgerly puts a five-line AGENTS.md inside migrations/ that explains how to create a new migration with Alembic. Agents working on fees never see it. Agents that go near migrations always do.

Some tools also let a file import other files by reference. Be careful: importing docs/gst.md into the root file puts all of it back into every session, which undoes the routing you just built. Keep routes as plain pointers the agent can follow when it needs them.

What instruction files cannot do

An instruction is a request. The model reads it and usually follows it. "Usually" is the problem. On Ledgerly, adding "Never edit merged migrations" to the instruction file cut migration edits from 6 in 20 runs to 2 in 20. That is a big improvement and still not zero. For a rule whose violation can corrupt a production database, 2 in 20 is not acceptable.

So each rule belongs at one of three strengths:

Instruction file only

  • Preferences and conventions
  • "Prefer small functions", "name tests after behaviour"
  • A miss costs a review comment

Instruction file plus gate

  • Things that must be true at the end
  • "Tests pass", "lint is clean", "behaviour change has a test"
  • A miss is caught before "done" is accepted

Instruction file plus permission

  • Things that must never happen at all
  • "No edits to merged migrations", "no reading .env", "no git push"
  • A miss is blocked before it runs

Keep the rule in the instruction file even when the harness enforces it. The instruction tells the agent why and what to do instead. The enforcement guarantees the outcome. An agent that knows the migration rule creates a new migration at the first attempt. An agent that only meets the permission block wastes turns discovering the rule by trial and error.

Check your understanding

0 of 3 answered

1.Ledgerly's old instruction file had 40 lines of GST rules. Only tax tasks need them. What should happen to them?

2.After adding "Never edit merged migrations" to the instruction file, agents still edited them in 2 of 20 runs. What is the right next step?

3.Which line passes the "check" test for an instruction file?