Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Model, harness, repo: who owns which failure
When an agent run goes wrong, the first reaction is usually to blame the model: "it got confused", "it hallucinated", "we need the bigger one". Sometimes that is true. Much more often, the model did a reasonable thing with bad inputs. It was given a tool that dumped 40,000 lines of output, or it was never told the test command, or it was asked to work in a file whose function names mean nothing.
Blaming the model is attractive because it asks nothing of you. You wait for the next release. Blaming the right layer is more useful, because two of the three layers are yours to fix this afternoon.
This lesson gives you a way to sort failures into three layers, a table of real Ledgerly failures sorted that way, and a habit — the failure log — that turns single bad runs into lasting fixes.
Three layers
Every coding-agent run has three layers, and each can fail on its own.
- The model turns a message list into the next action. It owns reasoning and knowledge: understanding the task, reading code correctly, writing correct Python.
- The harness is the program around the model. It owns what the model sees (instructions, tool results), what it can do (tools, permissions), when it stops (limits, gates), what it remembers (progress files) and what gets recorded (logs).
- The repository is the world the agent works in. It owns how easy it is to find things, how fast and trustworthy the tests are, and whether the rules a human needs are written down anywhere.
The layers overlap at the edges, but in practice most failures have one clear owner, and the owner decides the fix.
The contractor question
To find the owner, ask one question:
Would a strong human contractor, on their first day, with exactly the same instructions, the same tools and the same repository, have succeeded?
If the answer is no, the problem is not the model. A contractor told "add a late fee" with no test command, no warning about the flaky test and a shell that prints 40,000 lines per command would also struggle. That is a harness failure or a repo failure.
If the answer is yes — the contractor would have read the error message, seen the obvious bug and fixed it — then the model really did fall short. Now it is worth trying a stronger model, a different prompt, or breaking the task into smaller pieces.
The question works because it forces you to imagine the agent's actual situation. It is easy to forget that the agent never heard the hallway conversation where the team agreed not to touch utils.py.
Ledgerly failures, sorted
Here are eight failures from real agent runs on Ledgerly, sorted by the contractor question.
| Failure | Contractor succeeds? | Owner | Cheapest fix |
|---|---|---|---|
| Ran the full 41 s suite after every edit; task took 70 minutes | Yes, if told the fast command | Repo | Document pytest tests/unit -q (9 s) as the inner-loop command |
| Skipped the flaky PDF test to get green | Yes, if told the test is flaky | Repo and harness | Mark it flaky; gate forbids removing tests |
| Said "all tests pass" after running one file | No check existed | Harness | Completion gate runs the full checks |
Edited migrations/0019_add_gstin.py in place | Yes, if told migrations are append-only | Harness | Permission rule: writes under migrations/ need approval |
Read .env and copied the SMTP password into a test fixture | Maybe; nothing said not to | Harness | Deny reads of .env; run in a sandbox with fake secrets |
| Test output of 38,000 lines hid the one real error | Would scroll to the end | Harness | Tool keeps the head and tail of output, drops the middle |
Called utils.fx2() with arguments in the wrong order | Yes, after reading the function | Model (and repo) | Stronger model; also rename fx2 and add type hints |
| Misread GST rules and applied 18% twice on inter-state invoices | Maybe not — domain knowledge | Model | Put the rule and a worked example in the task |
Count the owners: of eight failures, only two are clearly model failures, and even one of those has a repo fix that helps humans too. This ratio is typical. In our experience across several repositories, somewhere between two thirds and three quarters of agent failures trace back to the harness or the repository.
Fixing the model layer
- Switch model, or wait for a new one
- Costs more per token, often
- Effect is broad but unpredictable
- You cannot test the change in advance
Fixing the harness or repo
- Add a rule, a check, a tool fix or a doc line
- Costs minutes of engineering
- Effect is narrow and certain
- You can write a test that proves it works
When it really is the model
Be honest about the other side. Some tasks need more capability than a model has. Signs that you are at the model's limit:
- The agent has everything it needs in its context (the right file, the error message, the instruction) and still draws the wrong conclusion, repeatedly, across several runs.
- A stronger model, given the identical harness and task, succeeds in most runs where the weaker one fails in most.
- The failures are in reasoning about the code, not in finding it, running it or knowing when to stop.
When you see this, change the model or shrink the task. Do not respond by piling on more harness rules. A 300-line instruction file cannot teach a model to reason about a tricky concurrency bug, and it will make every other run slower and more confused.
Keep a failure log
The habit that makes this lesson stick is simple: every time an agent run on your repository goes wrong, write one line in a shared file. Four fields are enough.
date task what went wrong owner fix2026-09-08 late fee on overdue skipped flaky PDF test to get green repo mark @flaky, gate blocks test removal2026-09-09 CSV export by date range ran full suite 31 times (21 min) repo AGENTS.md: inner-loop test command2026-09-11 round_money to Decimal claimed done after one file harness completion gate2026-09-12 GST inter-state split applied 18% twice model rule + worked example in task textAfter two weeks, sort by owner and count. The biggest group tells you where to spend your next day of work. On Ledgerly, the first two weeks produced 23 entries: 11 harness, 8 repo, 4 model. That count is the reason this course is mostly about the harness.
Check your understanding
0 of 3 answered
1.An agent keeps editing the wrong format_amount function, because Ledgerly has two: one in utils.py and one in invoices/pdf_helpers.py. Where does the fix belong first?
2.Which observation is the best evidence that you have hit the model's limit, not a harness gap?
3.Why write each failure into a shared log instead of just fixing it on the spot?