Harness Engineering: Making Coding Agents Dependable

Progress files the harness enforces


Ledgerly's first attempt at cross-session memory was simple. The instruction file said: "Keep PROGRESS.md up to date. Mark features as done when you finish them." For three sessions it looked great. The agent kept a tidy checklist, with dates and short notes.

Then the team looked closer. Feature F2 was marked "done". Its tests had never passed. The agent had written "done" in the file, then run the tests, then run out of turns while fixing them. The next session read PROGRESS.md, saw F2 done, and started F3. A half-working reminder email sat in the codebase for a week, trusted because a file said so.

The file was a memo: a record of what the agent believed. What the team needed was a gate: a record of what was verified, that nothing else could change.

How a feature's status is allowed to changetodo —writtenby a humanin_progress— set bythe harnessGate andverifycommand passdone — neverset by the agentA failed attempt stays in_progress; three failures mean blocked.
Because only a passed gate can mark a feature done, the next session is told only what was actually verified.

Memo versus gate

Progress memo

  • The agent writes it, whenever it likes
  • "Done" means the agent thinks it is done
  • Easy to start with: one line in the instructions
  • Drifts from reality the first time a session ends badly

Progress gate

  • The harness writes status; the agent only supplies notes
  • "Done" means the verification command passed
  • Needs a small amount of harness code
  • Stays true even when sessions fail, crash or run out of turns

The principle is the same as the completion gate: the agent's claims are inputs, not facts. The only difference is time. The completion gate checks a claim at the end of one session. The progress gate makes sure that what the next session is told is also true.

The file

Kite keeps its progress in kite-progress.json at the repository root. Here it is on Monday morning:

JSON
{  "goal": "Ledgerly: late fees on overdue invoices, shown everywhere a customer sees money.",  "features": [    {"id": "F1", "title": "2% late fee on invoices unpaid more than 30 days after the due date",     "status": "done", "verify": "python -m pytest -q tests/unit/test_late_fee.py",     "attempts": 1, "notes": "Fee is on the unpaid balance, computed when read. Added fees.apply_late_fee."},    {"id": "F2", "title": "Reminder emails show the late fee and the new total",     "status": "in_progress", "verify": "python -m pytest -q tests/unit/test_reminders.py",     "attempts": 1, "notes": "Template change is done. Patch ledgerly.mail.send, not smtplib.SMTP."},    {"id": "F3", "title": "CSV export of invoices for a date range", "status": "todo",     "verify": "python -m pytest -q tests/unit/test_export_csv.py", "notes": ""},    {"id": "Q4", "title": "round_money in ledgerly/utils.py uses float; round_money(2.675) returns 2.67, not 2.68",     "status": "proposed", "verify": "", "notes": "found during F1"}  ],  "log": ["2026-09-19 F1 attempt 1: done in 24 turns, 1148210 tokens",          "2026-09-19 F2 attempt 1: out_of_turns in 60 turns, 2410322 tokens"]}

Why JSON and not a Markdown checklist? Three reasons. The harness reads and writes it, so it must be machine-readable. Each feature carries a verify command, which is data, not prose. And in practice models are less inclined to casually rewrite a structured JSON file than a friendly Markdown list, so accidental edits are rarer. People can still read it easily.

Every feature has an acceptance command in verify. That is what makes the gate possible: "done" is defined by a command that passes, not by a sentence that sounds finished.

Who may change what

The gate works only if each field has a clear owner:

FieldWritten byWhen
features, title, verifyA human, or an initialisation session a human reviewsBefore feature work starts
statusThe harness onlyAt session start (in progress) and end (done, in progress, blocked)
attempts, logThe harness onlyAt session end
notesThe harness, copying the agent's final messageAt session end
New proposed itemsThe harness, from the agent's QUEUE: linesAt session end
proposed to todoA human onlyWhenever they review the queue

The agent never writes the file directly. In Kite, writes to kite-progress.json are in the permission layer's ask tier, so an unattended session cannot touch it. And even if a session did edit it, Kite loads the file at the start of the session and saves its own copy at the end, which overwrites the edit.

The transitions

  1. todo to in_progress — the harness picks the first open feature and marks it at the start of a session.
  2. in_progress to done — only when the completion gate passes, and the gate includes the feature's verify command.
  3. in_progress stays in_progress — the session ended any other way; the attempt count goes up and the agent's summary becomes the notes.
  4. in_progress to blocked — after three failed attempts; a human must look before any agent tries again.
  5. proposed to todo — only a human, after reading the item and writing a verify command for it.

Here is the code that enforces steps 2 to 4, from Kite's progress.py:

Python
def record(data: dict, feature: dict, outcome) -> None:    feature["attempts"] = feature.get("attempts", 0) + 1    if outcome.status == "done":        feature["status"] = "done"    else:        feature["status"] = "blocked" if feature["attempts"] >= 3 else "in_progress"    feature["notes"] = outcome.summary[-1500:]

The outcome comes from the loop, and its status is done only when the gate passed. There is no path by which the model's words alone can set a feature to done. The notes keep the last 1,500 characters of the summary, because the end of a message is where agents put conclusions and next steps.

Saving is atomic: Kite writes to a temporary file, then renames it over the real one. A crash halfway through a save leaves the old file intact, never half a JSON document.

Trade-offs

A progress gate asks for work up front. Someone has to split the goal into features small enough for one session — on Ledgerly, 15 to 30 turns each — and write a verification command for each. A feature like "improve the reminders" cannot be gated, because nothing can check it. That is a real cost, and it is also the point: if you cannot write the check, you are not ready to hand the work to an unattended agent.

For a one-off task in a single session, you do not need any of this. A progress file pays back when work spans three or more sessions, or when sessions run unattended overnight.

Check your understanding

0 of 3 answered

1.An agent's final message says "F3 is complete; all export tests pass", but the session ended with status out_of_turns. What does Kite record for F3?

2.Why does every feature in the progress file need a verify command?

3.A session somehow edits kite-progress.json and sets F3 to done. What happens at the end of the session?