Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Instruction files that actually bind
Ledgerly's first instruction file was 310 lines long. Someone pasted in the team wiki: the architecture overview, GST rules, the deploy process, a style guide and, on line 214, the command for running the fast tests. Agents still ran the 41-second full suite after every edit. The one line that would have saved half an hour per task was there. It was buried under 4,000 tokens of text that did not matter for the task at hand.
Every major coding agent reads an instruction file at the start of a session. Claude Code reads CLAUDE.md. Codex, Cursor and a growing list of other tools read AGENTS.md. The mechanism is the same everywhere: the file's text is placed into the context before your task. That makes it powerful, because the agent sees it in every session. It also makes it expensive, because every line competes for the model's attention with the task itself.
This lesson is about writing a file that the agent actually follows, and about knowing where an instruction file's power ends.
Root, route, check
A good instruction file does three jobs, and only three.
- Root: the few facts that every session needs. What the project is, where things live, the commands, and the handful of rules whose violation causes damage.
- Route: pointers to deeper documents that only some sessions need. "Before changing tax code, read
docs/gst.md." The agent reads that file only when the task goes there. - Check: every rule should be something the agent or the harness can verify. "Write clean code" cannot be checked. "
ruff check ledgerly testspasses" can.
Here is Ledgerly's rewritten root file. It is 26 lines, about 450 tokens.
# Ledgerly: instructions for coding agentsLedgerly is a Flask app for GST invoices and payment reminders (Python 3.11).Code is in ledgerly/, tests in tests/, database migrations in migrations/.## Commands- Fast tests while you work (9 s): python -m pytest tests/unit -q- A single test file: python -m pytest tests/unit/test_fees.py -q- Full check before you finish (45 s): bash scripts/check.sh- Lint only: ruff check ledgerly tests## Rules- Money is Decimal, never float. Amounts are rounded to 2 places, half-up.- Every behaviour change needs a test in tests/unit that fails without the change.- Never edit a migration that is already merged. Create a new one instead.- Do not open or print .env. Tests get settings from tests/fixtures/settings.py.- tests/integration/test_pdf_export.py is flaky (about 1 run in 8). It is marked flaky and excluded from check.sh. Do not skip, delete or "fix" it.- Work only on your task. Report unrelated problems in your final message, one per line, starting with QUEUE:.## Where to look- Fees and rounding: ledgerly/fees.py. Tax: ledgerly/tax.py, and read docs/gst.md first.- Invoice statuses and transitions: docs/invoice-states.md.- ledgerly/utils.py is legacy. Read only the function you need. Do not refactor it.- Migrations: read migrations/README.md before creating one.Look at what is not in it. No architecture essay: the agent can read the code. No list of libraries: pyproject.toml has that. No deploy process: the agent never deploys. No GST rules: they sit behind a route and cost nothing on tasks that do not touch tax.
Keep, move or delete
To shrink an old file, go through it line by line and give every line one of three verdicts. The test for each line is: if this line were missing, would a session go wrong in a way I can name?
| Line from the old 310-line file | Verdict | Why |
|---|---|---|
| "Write clean, maintainable code." | Delete | Not checkable, and the model already tries |
| "We use Flask 3 and SQLAlchemy 2." | Delete | Visible in pyproject.toml |
"Run fast tests with python -m pytest tests/unit -q." | Keep, move to the top | Saves about 30 minutes per task |
| "Never edit merged migrations." | Keep, and enforce in the harness | Too costly to leave to a request |
| 40 lines explaining GST rules | Move to docs/gst.md | Needed only by tax tasks |
| 25 lines on the deploy process | Delete | Agents never deploy Ledgerly |
| "Be careful with utils.py." | Rewrite | Say what to do: read one function, no refactors |
| "IMPORTANT: ALWAYS run ALL tests after EVERY change!!!" | Delete | Contradicts the fast-test rule; capitals do not help |
The last row is common and worth a word. Shouting in capital letters is a sign that an earlier rule was ignored. The usual reason it was ignored is that it conflicted with another rule or was buried. Fix the conflict or the position. Louder text just adds noise, and current models tend to over-apply emphasised rules to cases where they do not fit.
Two other signs that a line should go: it describes something that is already true of the code (the model will see it anyway), or nobody can remember which failure it was added to prevent.
Where instruction files live
Most tools support more than one instruction file. A root file loads at the start. Files in subdirectories load when the agent works in that folder, and for AGENTS.md the nearest file generally takes precedence. This is routing done by the tool for you.
Use it for rules that only matter in one place. Ledgerly puts a five-line AGENTS.md inside migrations/ that explains how to create a new migration with Alembic. Agents working on fees never see it. Agents that go near migrations always do.
Some tools also let a file import other files by reference. Be careful: importing docs/gst.md into the root file puts all of it back into every session, which undoes the routing you just built. Keep routes as plain pointers the agent can follow when it needs them.
What instruction files cannot do
An instruction is a request. The model reads it and usually follows it. "Usually" is the problem. On Ledgerly, adding "Never edit merged migrations" to the instruction file cut migration edits from 6 in 20 runs to 2 in 20. That is a big improvement and still not zero. For a rule whose violation can corrupt a production database, 2 in 20 is not acceptable.
So each rule belongs at one of three strengths:
Instruction file only
- Preferences and conventions
- "Prefer small functions", "name tests after behaviour"
- A miss costs a review comment
Instruction file plus gate
- Things that must be true at the end
- "Tests pass", "lint is clean", "behaviour change has a test"
- A miss is caught before "done" is accepted
Instruction file plus permission
- Things that must never happen at all
- "No edits to merged migrations", "no reading
.env", "nogit push" - A miss is blocked before it runs
Keep the rule in the instruction file even when the harness enforces it. The instruction tells the agent why and what to do instead. The enforcement guarantees the outcome. An agent that knows the migration rule creates a new migration at the first attempt. An agent that only meets the permission block wastes turns discovering the rule by trial and error.
Check your understanding
0 of 3 answered
1.Ledgerly's old instruction file had 40 lines of GST rules. Only tax tasks need them. What should happen to them?
2.After adding "Never edit merged migrations" to the instruction file, agents still edited them in 2 of 20 runs. What is the right next step?
3.Which line passes the "check" test for an instruction file?