Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Milestone 6: the event log, replay and a full run
Kite now works, remembers and verifies. It still cannot show you what happened inside a session. When F5 fails at 2 a.m., all you have is a status and a summary. This last milestone adds kite/events.py — the event log, the wrappers that write it, and replay — and then runs the finished harness on Ledgerly.
The loop does not change. The observability lesson explained why: logging is added by wrapping the model, the toolbox and the gate in objects with the same interface.
The log and its wrappers
1"""An append-only event log, wrappers that write to it, and replay from it."""2import json3import time4from pathlib import Path56from .model import Reply, ScriptedModel789class EventLog:10 def __init__(self, path: Path):11 self.path, self.start = path, time.monotonic()12 path.parent.mkdir(parents=True, exist_ok=True)13 (path.parent / ".gitignore").write_text("*\n") # logs never land in a commit1415 def emit(self, kind: str, **data) -> None:16 event = {"t": round(time.monotonic() - self.start, 2), "kind": kind, **data}17 with self.path.open("a") as f:18 f.write(json.dumps(event) + "\n")192021class LoggedModel:22 def __init__(self, inner, log: EventLog):23 self.inner, self.log, self.turn = inner, log, 02425 def complete(self, system, messages, tools):26 self.turn += 127 began = time.monotonic()28 reply = self.inner.complete(system, messages, tools)29 self.log.emit("model", turn=self.turn, ms=int((time.monotonic() - began) * 1000),30 stop=reply.stop_reason, content=reply.content,31 input_tokens=reply.input_tokens, output_tokens=reply.output_tokens)32 return replyemit opens the file, appends one line and closes it, every time. That is slower than keeping the file open, and it does not matter: a session writes a few hundred events, and each model call takes seconds. What it buys is that every finished event is on disk even if the process is killed mid-session. The .gitignore containing * makes the whole log folder invisible to git, so finish never commits logs.
Add the other two wrappers:
1class LoggedToolbox:2 def __init__(self, inner, log: EventLog):3 self.inner, self.log = inner, log45 def specs(self):6 return self.inner.specs()78 def run(self, call):9 began = time.monotonic()10 result = self.inner.run(call)11 self.log.emit("tool", name=call.name, args={k: str(v)[:300] for k, v in call.args.items()},12 ms=int((time.monotonic() - began) * 1000), is_error=result.is_error,13 output=result.output[:2000])14 return result151617class LoggedGate:18 def __init__(self, inner, log: EventLog):19 self.inner, self.log = inner, log2021 def check(self):22 verdict = self.inner.check()23 self.log.emit("gate", passed=verdict.passed, feedback=verdict.feedback[:3000])24 return verdictTool arguments and outputs are clipped in the log. The full write_file content is not lost: it is inside the model event's content blocks, because it was part of what the model said.
Reading the log back
1def read(path: Path) -> list[dict]:2 return [json.loads(line) for line in path.read_text().splitlines() if line.strip()]345def replay(path: Path) -> ScriptedModel:6 """A model that says, turn by turn, exactly what the recorded model said."""7 return ScriptedModel([Reply(e["content"], e["stop"], e["input_tokens"], e["output_tokens"])8 for e in read(path) if e["kind"] == "model"])91011def timeline(path: Path) -> str:12 rows = []13 for e in read(path):14 if e["kind"] == "model":15 rows.append(f"{e['t']:7.1f}s turn {e['turn']:<3} {e['stop']:<9} "16 f"in={e['input_tokens']:,} out={e['output_tokens']:,}")17 elif e["kind"] == "tool":18 flag = "ERROR" if e["is_error"] else "ok"19 rows.append(f"{e['t']:7.1f}s {flag:<5} {e['name']} {json.dumps(e['args'])[:60]}")20 elif e["kind"] == "gate":21 rows.append(f"{e['t']:7.1f}s gate {'PASSED' if e['passed'] else 'FAILED'}")22 return "\n".join(rows)replay is where the design decision from milestone 1 pays off. Because a Reply is just content blocks plus a stop reason, and the log stores exactly those, a log turns back into a ScriptedModel in one line. timeline prints the view you saw in the very first lesson.
The finished command line
Update main() in kite/__main__.py. Add import time and from .events import EventLog, LoggedGate, LoggedModel, LoggedToolbox, replay, timeline to the imports, add a --replay option after --unattended:
ap.add_argument("--replay", type=Path, help="replay model replies from an earlier run's log")and replace everything after feature["status"] = "in_progress" with:
1 name = f"{time.strftime('%Y%m%d-%H%M%S')}-{feature['id']}{'-replay' if args.replay else ''}"2 log = EventLog(root / cfg.log_dir / f"{name}.jsonl")3 model = replay(args.replay) if args.replay else AnthropicModel(cfg.model, cfg.max_tokens)4 policy = Policy(refuse if args.unattended else ask_human)5 toolbox = LoggedToolbox(Toolbox(root, cfg, check=policy.check), log)6 gate = LoggedGate(Gate(root, cfg.checks + [feature["verify"]], cfg.command_timeout), log)7 log.emit("session_start", feature=feature["id"], model=cfg.model, replay=str(args.replay or ""))8 outcome = run_session(progress.briefing(data, feature), LoggedModel(model, log), toolbox, cfg,9 system=SYSTEM, gate=gate)10 log.emit("session_end", status=outcome.status, turns=outcome.turns, tokens=outcome.tokens)11 print(timeline(log.path))12 if args.replay: # a replay is for looking, not for saving13 return 0 if outcome.status == "done" else 114 commit = progress.finish(root, path, data, feature, outcome)15 print(f"{feature['id']}: {outcome.status} in {outcome.turns} turns, "16 f"{outcome.tokens:,} tokens. Commit {commit}. Log {log.path}")17 return 0 if outcome.status == "done" else 1Each log is named with the start time and the feature id, plus -replay for replays, so a replay never writes into the log it is reading. A replay prints its timeline and exits without touching the progress file or git: it is for looking, not saving. That is the whole of Kite.
The end-to-end test
The last test runs the complete program — command line, progress, policy, tools, gate, log and commit — with a scripted model. Create tests/test_end_to_end.py:
1import subprocess23from kite import progress4from kite.__main__ import main5from kite.events import read6from kite.model import ScriptedModel, say, use78FEATURE = {"id": "F4", "title": "Fees round half-up", "status": "todo",9 "verify": "python -m pytest -q tests/test_rounding.py", "notes": ""}10FIXED = ("from decimal import ROUND_HALF_UP, Decimal\n\n\ndef add_fee(total, rate):\n"11 " fee = Decimal(str(total)) * Decimal(str(rate)) / 100\n"12 " return float(fee.quantize(Decimal('0.01'), rounding=ROUND_HALF_UP))\n")13TEST = ("from ledgerly.money import add_fee\n\n\ndef test_half_paisa_rounds_up():\n"14 " assert add_fee(133.75, 2) == 2.68\n")15SCRIPT = [use("read_file", "c1", path="ledgerly/money.py"),16 use("write_file", "c2", path="ledgerly/money.py", content=FIXED),17 say("Done: fees now round half-up."), # no test yet: the gate says no18 use("write_file", "c3", path="tests/test_rounding.py", content=TEST),19 use("run", "c4", command="python -m pytest -q"),20 say("Fees use Decimal half-up; added tests/test_rounding.py.\nQUEUE: fees.py rounds twice")]212223def run_kite(repo, monkeypatch, *extra) -> int:24 progress.save(repo / "kite-progress.json", {"goal": "Ledgerly", "log": [], "features": [dict(FEATURE)]})25 monkeypatch.setattr("kite.__main__.AnthropicModel", lambda *a: ScriptedModel(SCRIPT))26 return main([str(repo), "--unattended", *extra])272829def gates(log) -> list[bool]:30 return [e["passed"] for e in read(log) if e["kind"] == "gate"]313233def test_full_session(repo, monkeypatch):34 assert run_kite(repo, monkeypatch) == 035 data = progress.load(repo / "kite-progress.json")36 assert [f["status"] for f in data["features"]] == ["done", "proposed"]37 assert gates(next((repo / ".kite" / "runs").glob("*-F4.jsonl"))) == [False, True]383940def test_replay_reproduces_the_run(repo, monkeypatch):41 run_kite(repo, monkeypatch)42 log = next((repo / ".kite" / "runs").glob("*-F4.jsonl"))43 subprocess.run(["git", "reset", "-q", "--hard", "HEAD~1"], cwd=repo, check=True)44 assert run_kite(repo, monkeypatch, "--replay", str(log)) == 045 assert gates(next((repo / ".kite" / "runs").glob("*-F4-replay.jsonl"))) == [False, True]The script is the float-rounding fix from the scope-control lesson in miniature. add_fee(133.75, 2) is 2.675 rupees; float rounding gives 2.67 and Decimal half-up gives 2.68. The scripted agent fixes the code, claims it is done without a test, is refused by the gate, writes the test, and passes. The replay test resets the repository to before the session and replays the log: the gate refuses and then accepts at the same points. monkeypatch swaps the real model for the script, so no API key is needed. The twelve tests from the six milestones, plus the three from the testing lesson if you added them, run in about eight seconds.
The full run on Ledgerly
- Prepare — in Ledgerly,
git worktree add ../ledgerly-agent -b agent/late-fees; commitAGENTS.md,scripts/check.shand the reviewedkite-progress.jsonthere. - Run — from the Kite folder,
while python -m kite ../ledgerly-agent --unattended; do :; done, inside the container from the hard-limits lesson. - Read — run the cost script from the observability lesson over
../ledgerly-agent/.kite/runs, and the diagnosis script over any failed log. - Review — read
git log --oneline agent/late-fees, each commit's diff, and the proposed items in the progress file. - Decide — open a pull request for the done features; fix instructions or repo for the failures; promote or drop the proposals.
On Ledgerly's five late-fee features, the batch looked like this:
| Feature | Status | Turns | Tokens | Notes |
|---|---|---|---|---|
| F1 Late fee on overdue invoices | done | 24 | 1.15 M | Gate caught a missing test at turn 19 |
| F2 Fee shown in reminder emails | done | 9 | 0.31 M | Resumed from Friday's notes |
| F3 CSV export for a date range | done | 14 | 0.52 M | Used ledgerly.clock after the instruction fix |
| F4 Money rounding uses Decimal | done | 31 | 1.62 M | Promoted from proposal Q4 |
| F5 Late fee line on the PDF invoice | gate_failed | 27 | 1.38 M | PDF renderer missing in the container |
About 5 million tokens in total — roughly 25 USD at list input prices, and about 6 USD as billed, because most input was served from the prompt cache. Four features arrived as four clean, reviewed commits. F5's failure was not the agent's: its gate needed the headless PDF renderer, which the container did not have. The diagnosis script showed three identical gate failures naming the missing binary, and the fix was one line in the container image.
Check your understanding
0 of 3 answered
1.Why can Kite add the event log in this milestone without changing loop.py at all?
2.A replay runs against a fresh checkout of the starting commit. Which parts of the session come from the log, and which run live?
3.In the Ledgerly batch, F5 ended gate_failed because the container had no PDF renderer. Which layer owned that failure?