Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
What you will build: Kite on Ledgerly
You ask a coding agent to add a late fee to an invoicing app: an invoice still unpaid 30 days after its due date gets a 2% charge. Forty minutes later the agent replies: "Done. Added the late fee and all tests pass." You open the diff. The fee is there and it looks right. But the agent also rewrote round_money() in a shared utility file, edited an old database migration in place, and put a skip marker on a PDF test that "kept failing". And "all tests pass" means one test file — the one it ran last.
None of these edits is stupid on its own. Each one made sense from what the agent could see at that moment. The real problem is that nothing around the agent stopped it, checked its claim, or even recorded why it made each choice. That "something around the agent" is the subject of this course.
This lesson shows you the finished product first: a small harness called Kite, running against a small Flask app called Ledgerly. Every later lesson explains one piece of it, and the last section has you build all of it.
Ledgerly: the repository we break agents on
Ledgerly is a Python and Flask app that small businesses in India use to send GST invoices and chase payments. It is small enough to understand in an afternoon, and messy enough to be realistic.
| Fact | Value |
|---|---|
| Tracked files | 41 (33 Python modules, about 5,800 lines) |
| Stack | Python 3.11, Flask 3, SQLAlchemy, Alembic migrations, pytest, ruff |
| Test suite | 214 tests; unit tests take 9 s, the full suite 41 s |
| Flaky test | tests/integration/test_pdf_export.py fails about 1 run in 8 |
| Legacy file | ledgerly/utils.py: 640 lines, 31 functions, imported by 17 modules |
| Secrets | .env holds SMTP_PASSWORD and a staging DATABASE_URL |
Every row in that table caused a real agent failure. The flaky test invites the agent to "fix" it by skipping it. The legacy file invites a refactor nobody asked for. The .env file can be read by any process with file access, including the agent. And the 41-second suite, run after every small edit, turned a 20-minute task into a 70-minute one.
Ledgerly is not a bad codebase. It is an ordinary one. Most teams have a flaky test, a file everyone is afraid of, and secrets sitting next to the code. An agent that only works on clean repositories does not work.
Kite: the harness you will build
Kite is a minimal coding-agent harness: about 450 lines of Python in nine files, with one model SDK hidden behind a small interface. You point it at a repository and it works on one feature per session. A session on Ledgerly looks like this:
$ python -m kite ../ledgerly --unattended 0.0s turn 1 tool_use in=4,120 out=156 0.4s ok run {"command": "grep -rn due_date ledgerly"} 5.9s turn 2 tool_use in=6,890 out=88 6.0s ok read_file {"path": "ledgerly/invoices/service.py"} ... 402.7s turn 19 end_turn in=71,340 out=210 441.3s gate FAILED ... 530.4s gate PASSEDF1: done in 24 turns, 1,148,210 tokens. Commit 4c1e9a2. Log .kite/runs/20260921-101502-F1.jsonlRead it as a story. The agent searched, read files, edited, and at turn 19 said it was finished. Kite did not take its word. The completion gate ran the project's checks and found a problem: a source file had changed but no test had. The agent got that failure back as a message, wrote the missing test, and the gate passed at turn 24. Only then did Kite mark the feature done, make a git commit, and close the event log.
Here are the six parts of Kite, and the failure each one exists to prevent:
| Part | Kite file | The failure it prevents |
|---|---|---|
| Loop and message list | loop.py, model.py | Runs that never stop, or stop in a broken state |
| Tools | tools.py | Output the model cannot use: 50,000-line logs, crashes instead of errors |
| Permissions | permissions.py | git push, rm -rf, edits to migrations and .env |
| Completion gate | gate.py | "All tests pass" when they do not |
| Progress file | progress.py | Each new session starting blind, redoing or undoing earlier work |
| Event log and replay | events.py | Failures that nobody can explain or reproduce |
The other two files are small: config.py holds limits and the model id, and __main__.py wires the parts together.
The core idea of this course
Why put the weight on the harness and not on the model? Because of what you control. You cannot change a model's weights. You choose from a handful of models, and they improve on the vendor's schedule, not yours. The harness is code you own completely. You can test it, change it in ten minutes, and see the effect on the next run.
The numbers make this concrete. Moving to a better model might cut Ledgerly's failure rate on a task from 40% to 30%. Adding a completion gate turns every false "done" into a visible failure, whatever the model. The gate does not make the model smarter. It makes the model's mistakes impossible to miss.
This does not mean the model is unimportant. A harness cannot make a weak model do a hard refactor, and a harness with no model is just a script. The honest claim is narrower: once you have picked a capable model, most of the remaining reliability you can buy comes from the harness, and it is the cheapest place to buy it.
The course map
- The reliability gap — the three layers and who owns each failure, why long runs decay, the repository as the agent's whole world, and the loop at the centre of every agent.
- Contracts inside one session — instruction files that bind, scope control, hard limits, and completion gates that check the claim of "done".
- Memory across sessions — why every session starts from zero, progress files as a gate, a separate initialisation session, and git as save points.
- Observing and testing the harness — event logs and budgets, a checklist for reading a failed run, and tests that use a scripted fake model.
- Build it — six milestones that assemble Kite and end with a full run on Ledgerly.
You need Python 3.11 or newer, git, and an API key for the real runs. The tests in the build section use a scripted fake model, so you can build and test all of Kite without spending a rupee on tokens. Kite's real adapter uses the Anthropic Messages API, but the model sits behind a one-method interface, so you can swap in another provider without touching the rest.
Check your understanding
0 of 3 answered
1.An agent on Ledgerly reports "all tests pass", but it only ran tests/test_invoices.py. Which Kite part exists to catch this?
2.Your team wants more reliable agent runs. Why does this course put most of the effort into the harness rather than the model?
3.Which Ledgerly fact is most likely to tempt an agent into editing tests to make its work "pass"?