Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Milestone 5: the progress file and resuming
After milestone 4, one Kite session is reliable. Real work needs many sessions: a goal split into features, sessions that fail and must be resumed, and nights where Kite works through a list unattended. This milestone builds kite/progress.py, the memory that carries work from one session to the next, and changes the command line to work from it.
Everything here comes from the section on memory across sessions: a fresh session per feature, a briefing built from a file the harness owns, statuses that only verified outcomes can change, and a git commit as a save point after every session.
Loading, saving and choosing
1"""The progress file: the only memory that survives from one session to the next."""2import json3import subprocess4from datetime import date5from pathlib import Path678def load(path: Path) -> dict:9 return json.loads(path.read_text())101112def save(path: Path, data: dict) -> None:13 tmp = path.with_suffix(".tmp")14 tmp.write_text(json.dumps(data, indent=2) + "\n")15 tmp.replace(path) # atomic: a crash never leaves half a file161718def next_feature(data: dict) -> dict | None:19 return next((f for f in data["features"] if f["status"] in ("todo", "in_progress")), None)save writes a temporary file and renames it over the real one. On the same file system a rename is atomic: at every moment, the path holds either the complete old file or the complete new one. next_feature takes the first feature that is todo or in_progress, so an unfinished feature is always resumed before a new one starts, and blocked and proposed items are skipped until a human acts.
Briefing in, record out
1def briefing(data: dict, feature: dict) -> str:2 done = [f"- {f['id']} {f['title']}" for f in data["features"] if f["status"] == "done"]3 return "\n".join([4 f"Project: {data['goal']}",5 "Already done and verified by the harness (do not redo):", *(done or ["- nothing yet"]),6 "",7 f"Your task this session: {feature['id']} - {feature['title']}",8 f"It is accepted when this passes: {feature['verify']}, and so do the standard checks.",9 f"Notes from earlier attempts: {feature.get('notes') or 'none - this is the first attempt'}",10 ])111213def record(data: dict, feature: dict, outcome) -> None:14 feature["attempts"] = feature.get("attempts", 0) + 115 if outcome.status == "done":16 feature["status"] = "done"17 else:18 feature["status"] = "blocked" if feature["attempts"] >= 3 else "in_progress"19 feature["notes"] = outcome.summary[-1500:]20 for line in outcome.summary.splitlines(): # scope control: queue it, never do it21 if line.startswith("QUEUE:"):22 data["features"].append({"id": f"Q{len(data['features']) + 1}", "title": line[6:].strip(),23 "status": "proposed", "verify": "", "notes": f"found during {feature['id']}"})24 data["log"].append(f"{date.today()} {feature['id']} attempt {feature['attempts']}: "25 f"{outcome.status} in {outcome.turns} turns, {outcome.tokens} tokens")briefing is the handoff from the lesson on why sessions start from zero: goal, verified work, this session's task and its acceptance command, and notes from earlier attempts. It becomes the session's first user message, so the task is in view from turn 1.
record is the gate from the progress-file lesson. Only an outcome with status done — which the loop produces only when the completion gate passed — can mark a feature done. The agent's summary becomes the notes; QUEUE: lines become proposals a human must promote.
Save points
1def finish(root: Path, path: Path, data: dict, feature: dict, outcome) -> str:2 """Save the session's result and leave a clean tree; return the new commit hash."""3 if outcome.status != "done": # keep failed work, but out of the way4 subprocess.run(["git", "stash", "push", "-u", "-q", "-m", f"kite {feature['id']} failed"], cwd=root)5 record(data, feature, outcome)6 save(path, data)7 subprocess.run(["git", "add", "-A"], cwd=root, check=True)8 subprocess.run(["git", "commit", "-q", "-m", f"kite: {feature['id']} {outcome.status}"], cwd=root, check=True)9 return subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=root,10 capture_output=True, text=True).stdout.strip()This is the function from the lesson on git save points. A done feature becomes one commit with its code, tests and progress update. A failed attempt goes into a labelled stash, and only the progress update is committed. Either way the next session starts from a clean tree, which the gate's "what changed" rule depends on.
The command line, driven by the progress file
Replace main() in kite/__main__.py, and add from . import progress to the imports (Gate is already imported since milestone 4). The --task option goes away: the task now comes from the progress file, so update the module docstring to say python -m kite REPO [--unattended]: one session on the next open feature.
1def main(argv=None) -> int:2 ap = argparse.ArgumentParser(prog="kite")3 ap.add_argument("repo", type=Path)4 ap.add_argument("--unattended", action="store_true", help="refuse anything that needs a human")5 args = ap.parse_args(argv)6 cfg, root = Config(), args.repo.resolve()7 path = root / cfg.progress_file8 data = progress.load(path)9 feature = progress.next_feature(data)10 if feature is None:11 print(f"No open features in {path}.")12 return 213 feature["status"] = "in_progress"1415 policy = Policy(refuse if args.unattended else ask_human)16 toolbox = Toolbox(root, cfg, check=policy.check)17 gate = Gate(root, cfg.checks + [feature["verify"]], cfg.command_timeout)18 model = AnthropicModel(cfg.model, cfg.max_tokens)19 outcome = run_session(progress.briefing(data, feature), model, toolbox, cfg,20 system=SYSTEM, gate=gate)21 commit = progress.finish(root, path, data, feature, outcome)22 print(f"{feature['id']}: {outcome.status} in {outcome.turns} turns, "23 f"{outcome.tokens:,} tokens. Commit {commit}.")24 return 0 if outcome.status == "done" else 1The gate now includes the feature's own verify command after the standard checks. The exit code tells a script what happened: 0 for a finished feature, 1 for a failed session, 2 when nothing is left to do. That makes a batch a one-line shell loop:
while python -m kite ../ledgerly-agent --unattended; do :; doneIt keeps going while features finish and stops at the first failure or when the list is empty. Stopping at the first failure is a choice: a human looks at the problem before more sessions build on top of it. If you would rather let later, independent features continue, loop until the exit code is 2 and let blocked statuses collect the failures for the morning.
Features that fit in one session
The progress file only works if each feature fits comfortably in one session. On Ledgerly, a good feature takes 15 to 30 turns, changes one behaviour, and has one verify command that proves it. Compare:
| Too big | Right size |
|---|---|
| "Late fees everywhere" | "2% late fee on invoices unpaid more than 30 days after the due date" |
| "Improve reminders" | "Reminder emails show the late fee and the new total" |
| "Export and import invoices as CSV" | "CSV export of invoices for a date range" (import is a separate feature) |
A feature that is too big shows up in the log as out_of_turns with a summary that says "partially done". Split it and move on; a third attempt at a feature that is too big rarely succeeds. A feature that is too small — "add a docstring to apply_late_fee" — wastes the fixed cost of a session: the briefing, the reading, and a 45-second gate. Group small, related changes into one feature with one test.
Test it
Create tests/test_progress.py:
1import json2import subprocess34from kite import progress5from kite.loop import Outcome67DATA = {"goal": "Ledgerly late fees", "log": [], "features": [8 {"id": "F1", "title": "2% late fee", "status": "done", "verify": "true", "notes": ""},9 {"id": "F2", "title": "Fee on reminders", "status": "todo", "verify": "true", "notes": ""}]}101112def test_briefing_and_failed_attempt(repo):13 path = repo / "kite-progress.json"14 progress.save(path, json.loads(json.dumps(DATA)))15 subprocess.run("git add -A && git commit -qm progress", shell=True, cwd=repo, check=True)16 data = progress.load(path)17 feature = progress.next_feature(data)18 assert feature["id"] == "F2" and "- F1 2% late fee" in progress.briefing(data, feature)1920 (repo / "ledgerly" / "half_done.py").write_text("x = 1\n")21 outcome = Outcome("gate_failed", 12, 90_000, "Tests fail.\nQUEUE: utils.parse_date ignores time zones")22 progress.finish(repo, path, data, feature, outcome)23 saved = progress.load(path)24 assert saved["features"][1]["status"] == "in_progress"25 assert saved["features"][2]["status"] == "proposed"26 assert not (repo / "ledgerly" / "half_done.py").exists() # stashed, not lost27 assert "kite F2 failed" in subprocess.run(["git", "stash", "list"], cwd=repo,28 capture_output=True, text=True).stdoutjson.loads(json.dumps(DATA)) makes a deep copy, so the test never changes the module-level data. The test walks one full failed session: choose F2, brief it, fail, stash the half-done file, keep F2 in progress, turn the QUEUE: line into a proposal. For a real run, write Ledgerly's kite-progress.json by hand or with an initialisation session, commit it, and start the loop.
Check your understanding
0 of 3 answered
1.Why does next_feature pick in_progress features as well as todo ones?
2.A session ends out_of_tokens after editing three files. What does finish leave behind?
3.Why does Kite's batch loop while python -m kite ... --unattended; do :; done stop at the first failed session?