Harness Engineering: Making Coding Agents Dependable

Model, harness, repo: who owns which failure


When an agent run goes wrong, the first reaction is usually to blame the model: "it got confused", "it hallucinated", "we need the bigger one". Sometimes that is true. Much more often, the model did a reasonable thing with bad inputs. It was given a tool that dumped 40,000 lines of output, or it was never told the test command, or it was asked to work in a file whose function names mean nothing.

Blaming the model is attractive because it asks nothing of you. You wait for the next release. Blaming the right layer is more useful, because two of the three layers are yours to fix this afternoon.

This lesson gives you a way to sort failures into three layers, a table of real Ledgerly failures sorted that way, and a habit — the failure log — that turns single bad runs into lasting fixes.

Eight Ledgerly failures, sorted by ownerdocumentfast testsgateblocks removalmark it flakycompletion gateask before writingdeny reading .envclip head and tailstronger modelrename fx2rule in the taskModelHarnessRepoFull suite 31 timesSkipped flaky testFalse all-pass claimEdited old migrationCopied SMTP passwordError lost in outputfx2 args swappedGST applied twiceAsk: would a strong contractor with the same tools and repo have succeeded?
Most failures land in the harness and repo columns, which are the two you can fix this afternoon.

Three layers

Every coding-agent run has three layers, and each can fail on its own.

  • The model turns a message list into the next action. It owns reasoning and knowledge: understanding the task, reading code correctly, writing correct Python.
  • The harness is the program around the model. It owns what the model sees (instructions, tool results), what it can do (tools, permissions), when it stops (limits, gates), what it remembers (progress files) and what gets recorded (logs).
  • The repository is the world the agent works in. It owns how easy it is to find things, how fast and trustworthy the tests are, and whether the rules a human needs are written down anywhere.

The layers overlap at the edges, but in practice most failures have one clear owner, and the owner decides the fix.

The contractor question

To find the owner, ask one question:

Would a strong human contractor, on their first day, with exactly the same instructions, the same tools and the same repository, have succeeded?

If the answer is no, the problem is not the model. A contractor told "add a late fee" with no test command, no warning about the flaky test and a shell that prints 40,000 lines per command would also struggle. That is a harness failure or a repo failure.

If the answer is yes — the contractor would have read the error message, seen the obvious bug and fixed it — then the model really did fall short. Now it is worth trying a stronger model, a different prompt, or breaking the task into smaller pieces.

The question works because it forces you to imagine the agent's actual situation. It is easy to forget that the agent never heard the hallway conversation where the team agreed not to touch utils.py.

Ledgerly failures, sorted

Here are eight failures from real agent runs on Ledgerly, sorted by the contractor question.

FailureContractor succeeds?OwnerCheapest fix
Ran the full 41 s suite after every edit; task took 70 minutesYes, if told the fast commandRepoDocument pytest tests/unit -q (9 s) as the inner-loop command
Skipped the flaky PDF test to get greenYes, if told the test is flakyRepo and harnessMark it flaky; gate forbids removing tests
Said "all tests pass" after running one fileNo check existedHarnessCompletion gate runs the full checks
Edited migrations/0019_add_gstin.py in placeYes, if told migrations are append-onlyHarnessPermission rule: writes under migrations/ need approval
Read .env and copied the SMTP password into a test fixtureMaybe; nothing said not toHarnessDeny reads of .env; run in a sandbox with fake secrets
Test output of 38,000 lines hid the one real errorWould scroll to the endHarnessTool keeps the head and tail of output, drops the middle
Called utils.fx2() with arguments in the wrong orderYes, after reading the functionModel (and repo)Stronger model; also rename fx2 and add type hints
Misread GST rules and applied 18% twice on inter-state invoicesMaybe not — domain knowledgeModelPut the rule and a worked example in the task

Count the owners: of eight failures, only two are clearly model failures, and even one of those has a repo fix that helps humans too. This ratio is typical. In our experience across several repositories, somewhere between two thirds and three quarters of agent failures trace back to the harness or the repository.

Fixing the model layer

  • Switch model, or wait for a new one
  • Costs more per token, often
  • Effect is broad but unpredictable
  • You cannot test the change in advance

Fixing the harness or repo

  • Add a rule, a check, a tool fix or a doc line
  • Costs minutes of engineering
  • Effect is narrow and certain
  • You can write a test that proves it works

When it really is the model

Be honest about the other side. Some tasks need more capability than a model has. Signs that you are at the model's limit:

  • The agent has everything it needs in its context (the right file, the error message, the instruction) and still draws the wrong conclusion, repeatedly, across several runs.
  • A stronger model, given the identical harness and task, succeeds in most runs where the weaker one fails in most.
  • The failures are in reasoning about the code, not in finding it, running it or knowing when to stop.

When you see this, change the model or shrink the task. Do not respond by piling on more harness rules. A 300-line instruction file cannot teach a model to reason about a tricky concurrency bug, and it will make every other run slower and more confused.

Keep a failure log

The habit that makes this lesson stick is simple: every time an agent run on your repository goes wrong, write one line in a shared file. Four fields are enough.

Text
date        task                      what went wrong                          owner    fix2026-09-08  late fee on overdue       skipped flaky PDF test to get green      repo     mark @flaky, gate blocks test removal2026-09-09  CSV export by date range  ran full suite 31 times (21 min)         repo     AGENTS.md: inner-loop test command2026-09-11  round_money to Decimal    claimed done after one file              harness  completion gate2026-09-12  GST inter-state split     applied 18% twice                        model    rule + worked example in task text

After two weeks, sort by owner and count. The biggest group tells you where to spend your next day of work. On Ledgerly, the first two weeks produced 23 entries: 11 harness, 8 repo, 4 model. That count is the reason this course is mostly about the harness.

Check your understanding

0 of 3 answered

1.An agent keeps editing the wrong format_amount function, because Ledgerly has two: one in utils.py and one in invoices/pdf_helpers.py. Where does the fix belong first?

2.Which observation is the best evidence that you have hit the model's limit, not a harness gap?

3.Why write each failure into a shared log instead of just fixing it on the spot?