Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Progress files the harness enforces
Ledgerly's first attempt at cross-session memory was simple. The instruction file said: "Keep PROGRESS.md up to date. Mark features as done when you finish them." For three sessions it looked great. The agent kept a tidy checklist, with dates and short notes.
Then the team looked closer. Feature F2 was marked "done". Its tests had never passed. The agent had written "done" in the file, then run the tests, then run out of turns while fixing them. The next session read PROGRESS.md, saw F2 done, and started F3. A half-working reminder email sat in the codebase for a week, trusted because a file said so.
The file was a memo: a record of what the agent believed. What the team needed was a gate: a record of what was verified, that nothing else could change.
Memo versus gate
Progress memo
- The agent writes it, whenever it likes
- "Done" means the agent thinks it is done
- Easy to start with: one line in the instructions
- Drifts from reality the first time a session ends badly
Progress gate
- The harness writes status; the agent only supplies notes
- "Done" means the verification command passed
- Needs a small amount of harness code
- Stays true even when sessions fail, crash or run out of turns
The principle is the same as the completion gate: the agent's claims are inputs, not facts. The only difference is time. The completion gate checks a claim at the end of one session. The progress gate makes sure that what the next session is told is also true.
The file
Kite keeps its progress in kite-progress.json at the repository root. Here it is on Monday morning:
1{2 "goal": "Ledgerly: late fees on overdue invoices, shown everywhere a customer sees money.",3 "features": [4 {"id": "F1", "title": "2% late fee on invoices unpaid more than 30 days after the due date",5 "status": "done", "verify": "python -m pytest -q tests/unit/test_late_fee.py",6 "attempts": 1, "notes": "Fee is on the unpaid balance, computed when read. Added fees.apply_late_fee."},7 {"id": "F2", "title": "Reminder emails show the late fee and the new total",8 "status": "in_progress", "verify": "python -m pytest -q tests/unit/test_reminders.py",9 "attempts": 1, "notes": "Template change is done. Patch ledgerly.mail.send, not smtplib.SMTP."},10 {"id": "F3", "title": "CSV export of invoices for a date range", "status": "todo",11 "verify": "python -m pytest -q tests/unit/test_export_csv.py", "notes": ""},12 {"id": "Q4", "title": "round_money in ledgerly/utils.py uses float; round_money(2.675) returns 2.67, not 2.68",13 "status": "proposed", "verify": "", "notes": "found during F1"}14 ],15 "log": ["2026-09-19 F1 attempt 1: done in 24 turns, 1148210 tokens",16 "2026-09-19 F2 attempt 1: out_of_turns in 60 turns, 2410322 tokens"]17}Why JSON and not a Markdown checklist? Three reasons. The harness reads and writes it, so it must be machine-readable. Each feature carries a verify command, which is data, not prose. And in practice models are less inclined to casually rewrite a structured JSON file than a friendly Markdown list, so accidental edits are rarer. People can still read it easily.
Every feature has an acceptance command in verify. That is what makes the gate possible: "done" is defined by a command that passes, not by a sentence that sounds finished.
Who may change what
The gate works only if each field has a clear owner:
| Field | Written by | When |
|---|---|---|
features, title, verify | A human, or an initialisation session a human reviews | Before feature work starts |
status | The harness only | At session start (in progress) and end (done, in progress, blocked) |
attempts, log | The harness only | At session end |
notes | The harness, copying the agent's final message | At session end |
New proposed items | The harness, from the agent's QUEUE: lines | At session end |
proposed to todo | A human only | Whenever they review the queue |
The agent never writes the file directly. In Kite, writes to kite-progress.json are in the permission layer's ask tier, so an unattended session cannot touch it. And even if a session did edit it, Kite loads the file at the start of the session and saves its own copy at the end, which overwrites the edit.
The transitions
- todo to in_progress — the harness picks the first open feature and marks it at the start of a session.
- in_progress to done — only when the completion gate passes, and the gate includes the feature's
verifycommand. - in_progress stays in_progress — the session ended any other way; the attempt count goes up and the agent's summary becomes the notes.
- in_progress to blocked — after three failed attempts; a human must look before any agent tries again.
- proposed to todo — only a human, after reading the item and writing a
verifycommand for it.
Here is the code that enforces steps 2 to 4, from Kite's progress.py:
1def record(data: dict, feature: dict, outcome) -> None:2 feature["attempts"] = feature.get("attempts", 0) + 13 if outcome.status == "done":4 feature["status"] = "done"5 else:6 feature["status"] = "blocked" if feature["attempts"] >= 3 else "in_progress"7 feature["notes"] = outcome.summary[-1500:]The outcome comes from the loop, and its status is done only when the gate passed. There is no path by which the model's words alone can set a feature to done. The notes keep the last 1,500 characters of the summary, because the end of a message is where agents put conclusions and next steps.
Saving is atomic: Kite writes to a temporary file, then renames it over the real one. A crash halfway through a save leaves the old file intact, never half a JSON document.
Trade-offs
A progress gate asks for work up front. Someone has to split the goal into features small enough for one session — on Ledgerly, 15 to 30 turns each — and write a verification command for each. A feature like "improve the reminders" cannot be gated, because nothing can check it. That is a real cost, and it is also the point: if you cannot write the check, you are not ready to hand the work to an unattended agent.
For a one-off task in a single session, you do not need any of this. A progress file pays back when work spans three or more sessions, or when sessions run unattended overnight.
Check your understanding
0 of 3 answered
1.An agent's final message says "F3 is complete; all export tests pass", but the session ended with status out_of_turns. What does Kite record for F3?
2.Why does every feature in the progress file need a verify command?
3.A session somehow edits kite-progress.json and sets F3 to done. What happens at the end of the session?