Harness Engineering: Making Coding Agents Dependable

Milestone 6: the event log, replay and a full run


Kite now works, remembers and verifies. It still cannot show you what happened inside a session. When F5 fails at 2 a.m., all you have is a status and a summary. This last milestone adds kite/events.py — the event log, the wrappers that write it, and replay — and then runs the finished harness on Ledgerly.

The loop does not change. The observability lesson explained why: logging is added by wrapping the model, the toolbox and the gate in objects with the same interface.

The late-fee batch on Ledgerlydone241.15 Mdone90.31 Mdone140.52 Mdone311.62 Mgate_failed271.38 MStatusTurnsTokensF1 late feeF2 remindersF3 CSV exportF4 DecimalF5 PDF lineF5 failed because the container lacked the PDF renderer.
Same model as the unguarded run: four clean commits, no accepted false done, and the one failure explained in minutes.

The log and its wrappers

Python
"""An append-only event log, wrappers that write to it, and replay from it."""import jsonimport timefrom pathlib import Pathfrom .model import Reply, ScriptedModelclass EventLog:    def __init__(self, path: Path):        self.path, self.start = path, time.monotonic()        path.parent.mkdir(parents=True, exist_ok=True)        (path.parent / ".gitignore").write_text("*\n")   # logs never land in a commit    def emit(self, kind: str, **data) -> None:        event = {"t": round(time.monotonic() - self.start, 2), "kind": kind, **data}        with self.path.open("a") as f:            f.write(json.dumps(event) + "\n")class LoggedModel:    def __init__(self, inner, log: EventLog):        self.inner, self.log, self.turn = inner, log, 0    def complete(self, system, messages, tools):        self.turn += 1        began = time.monotonic()        reply = self.inner.complete(system, messages, tools)        self.log.emit("model", turn=self.turn, ms=int((time.monotonic() - began) * 1000),                      stop=reply.stop_reason, content=reply.content,                      input_tokens=reply.input_tokens, output_tokens=reply.output_tokens)        return reply

emit opens the file, appends one line and closes it, every time. That is slower than keeping the file open, and it does not matter: a session writes a few hundred events, and each model call takes seconds. What it buys is that every finished event is on disk even if the process is killed mid-session. The .gitignore containing * makes the whole log folder invisible to git, so finish never commits logs.

Add the other two wrappers:

Python
class LoggedToolbox:    def __init__(self, inner, log: EventLog):        self.inner, self.log = inner, log    def specs(self):        return self.inner.specs()    def run(self, call):        began = time.monotonic()        result = self.inner.run(call)        self.log.emit("tool", name=call.name, args={k: str(v)[:300] for k, v in call.args.items()},                      ms=int((time.monotonic() - began) * 1000), is_error=result.is_error,                      output=result.output[:2000])        return resultclass LoggedGate:    def __init__(self, inner, log: EventLog):        self.inner, self.log = inner, log    def check(self):        verdict = self.inner.check()        self.log.emit("gate", passed=verdict.passed, feedback=verdict.feedback[:3000])        return verdict

Tool arguments and outputs are clipped in the log. The full write_file content is not lost: it is inside the model event's content blocks, because it was part of what the model said.

Reading the log back

Python
def read(path: Path) -> list[dict]:    return [json.loads(line) for line in path.read_text().splitlines() if line.strip()]def replay(path: Path) -> ScriptedModel:    """A model that says, turn by turn, exactly what the recorded model said."""    return ScriptedModel([Reply(e["content"], e["stop"], e["input_tokens"], e["output_tokens"])                          for e in read(path) if e["kind"] == "model"])def timeline(path: Path) -> str:    rows = []    for e in read(path):        if e["kind"] == "model":            rows.append(f"{e['t']:7.1f}s  turn {e['turn']:<3} {e['stop']:<9} "                        f"in={e['input_tokens']:,} out={e['output_tokens']:,}")        elif e["kind"] == "tool":            flag = "ERROR" if e["is_error"] else "ok"            rows.append(f"{e['t']:7.1f}s    {flag:<5} {e['name']} {json.dumps(e['args'])[:60]}")        elif e["kind"] == "gate":            rows.append(f"{e['t']:7.1f}s  gate {'PASSED' if e['passed'] else 'FAILED'}")    return "\n".join(rows)

replay is where the design decision from milestone 1 pays off. Because a Reply is just content blocks plus a stop reason, and the log stores exactly those, a log turns back into a ScriptedModel in one line. timeline prints the view you saw in the very first lesson.

The finished command line

Update main() in kite/__main__.py. Add import time and from .events import EventLog, LoggedGate, LoggedModel, LoggedToolbox, replay, timeline to the imports, add a --replay option after --unattended:

Python
    ap.add_argument("--replay", type=Path, help="replay model replies from an earlier run's log")

and replace everything after feature["status"] = "in_progress" with:

Python
    name = f"{time.strftime('%Y%m%d-%H%M%S')}-{feature['id']}{'-replay' if args.replay else ''}"    log = EventLog(root / cfg.log_dir / f"{name}.jsonl")    model = replay(args.replay) if args.replay else AnthropicModel(cfg.model, cfg.max_tokens)    policy = Policy(refuse if args.unattended else ask_human)    toolbox = LoggedToolbox(Toolbox(root, cfg, check=policy.check), log)    gate = LoggedGate(Gate(root, cfg.checks + [feature["verify"]], cfg.command_timeout), log)    log.emit("session_start", feature=feature["id"], model=cfg.model, replay=str(args.replay or ""))    outcome = run_session(progress.briefing(data, feature), LoggedModel(model, log), toolbox, cfg,                          system=SYSTEM, gate=gate)    log.emit("session_end", status=outcome.status, turns=outcome.turns, tokens=outcome.tokens)    print(timeline(log.path))    if args.replay:                        # a replay is for looking, not for saving        return 0 if outcome.status == "done" else 1    commit = progress.finish(root, path, data, feature, outcome)    print(f"{feature['id']}: {outcome.status} in {outcome.turns} turns, "          f"{outcome.tokens:,} tokens. Commit {commit}. Log {log.path}")    return 0 if outcome.status == "done" else 1

Each log is named with the start time and the feature id, plus -replay for replays, so a replay never writes into the log it is reading. A replay prints its timeline and exits without touching the progress file or git: it is for looking, not saving. That is the whole of Kite.

The end-to-end test

The last test runs the complete program — command line, progress, policy, tools, gate, log and commit — with a scripted model. Create tests/test_end_to_end.py:

Python
import subprocessfrom kite import progressfrom kite.__main__ import mainfrom kite.events import readfrom kite.model import ScriptedModel, say, useFEATURE = {"id": "F4", "title": "Fees round half-up", "status": "todo",           "verify": "python -m pytest -q tests/test_rounding.py", "notes": ""}FIXED = ("from decimal import ROUND_HALF_UP, Decimal\n\n\ndef add_fee(total, rate):\n"         "    fee = Decimal(str(total)) * Decimal(str(rate)) / 100\n"         "    return float(fee.quantize(Decimal('0.01'), rounding=ROUND_HALF_UP))\n")TEST = ("from ledgerly.money import add_fee\n\n\ndef test_half_paisa_rounds_up():\n"        "    assert add_fee(133.75, 2) == 2.68\n")SCRIPT = [use("read_file", "c1", path="ledgerly/money.py"),          use("write_file", "c2", path="ledgerly/money.py", content=FIXED),          say("Done: fees now round half-up."),                 # no test yet: the gate says no          use("write_file", "c3", path="tests/test_rounding.py", content=TEST),          use("run", "c4", command="python -m pytest -q"),          say("Fees use Decimal half-up; added tests/test_rounding.py.\nQUEUE: fees.py rounds twice")]def run_kite(repo, monkeypatch, *extra) -> int:    progress.save(repo / "kite-progress.json", {"goal": "Ledgerly", "log": [], "features": [dict(FEATURE)]})    monkeypatch.setattr("kite.__main__.AnthropicModel", lambda *a: ScriptedModel(SCRIPT))    return main([str(repo), "--unattended", *extra])def gates(log) -> list[bool]:    return [e["passed"] for e in read(log) if e["kind"] == "gate"]def test_full_session(repo, monkeypatch):    assert run_kite(repo, monkeypatch) == 0    data = progress.load(repo / "kite-progress.json")    assert [f["status"] for f in data["features"]] == ["done", "proposed"]    assert gates(next((repo / ".kite" / "runs").glob("*-F4.jsonl"))) == [False, True]def test_replay_reproduces_the_run(repo, monkeypatch):    run_kite(repo, monkeypatch)    log = next((repo / ".kite" / "runs").glob("*-F4.jsonl"))    subprocess.run(["git", "reset", "-q", "--hard", "HEAD~1"], cwd=repo, check=True)    assert run_kite(repo, monkeypatch, "--replay", str(log)) == 0    assert gates(next((repo / ".kite" / "runs").glob("*-F4-replay.jsonl"))) == [False, True]

The script is the float-rounding fix from the scope-control lesson in miniature. add_fee(133.75, 2) is 2.675 rupees; float rounding gives 2.67 and Decimal half-up gives 2.68. The scripted agent fixes the code, claims it is done without a test, is refused by the gate, writes the test, and passes. The replay test resets the repository to before the session and replays the log: the gate refuses and then accepts at the same points. monkeypatch swaps the real model for the script, so no API key is needed. The twelve tests from the six milestones, plus the three from the testing lesson if you added them, run in about eight seconds.

The full run on Ledgerly

  1. Prepare — in Ledgerly, git worktree add ../ledgerly-agent -b agent/late-fees; commit AGENTS.md, scripts/check.sh and the reviewed kite-progress.json there.
  2. Run — from the Kite folder, while python -m kite ../ledgerly-agent --unattended; do :; done, inside the container from the hard-limits lesson.
  3. Read — run the cost script from the observability lesson over ../ledgerly-agent/.kite/runs, and the diagnosis script over any failed log.
  4. Review — read git log --oneline agent/late-fees, each commit's diff, and the proposed items in the progress file.
  5. Decide — open a pull request for the done features; fix instructions or repo for the failures; promote or drop the proposals.

On Ledgerly's five late-fee features, the batch looked like this:

FeatureStatusTurnsTokensNotes
F1 Late fee on overdue invoicesdone241.15 MGate caught a missing test at turn 19
F2 Fee shown in reminder emailsdone90.31 MResumed from Friday's notes
F3 CSV export for a date rangedone140.52 MUsed ledgerly.clock after the instruction fix
F4 Money rounding uses Decimaldone311.62 MPromoted from proposal Q4
F5 Late fee line on the PDF invoicegate_failed271.38 MPDF renderer missing in the container

About 5 million tokens in total — roughly 25 USD at list input prices, and about 6 USD as billed, because most input was served from the prompt cache. Four features arrived as four clean, reviewed commits. F5's failure was not the agent's: its gate needed the headless PDF renderer, which the container did not have. The diagnosis script showed three identical gate failures naming the missing binary, and the fix was one line in the container image.

Check your understanding

0 of 3 answered

1.Why can Kite add the event log in this milestone without changing loop.py at all?

2.A replay runs against a fresh checkout of the starting commit. Which parts of the session come from the log, and which run live?

3.In the Ledgerly batch, F5 ended gate_failed because the container had no PDF renderer. Which layer owned that failure?