Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Testing the harness itself
The Ledgerly team once tightened Kite's gate with a new rule: a changed source file must come with a changed test file. The change looked obviously right, and it was merged in an afternoon. That night, a session that wrote no test at all passed the gate. Running pytest had created new tests/__pycache__/*.pyc files, and the rule counted them as "changed tests". A check designed to catch missing tests was fooled by the side effect of running them.
The harness is code. It has bugs like any code, and its bugs are expensive, because a wrong gate or a wrong permission rule quietly affects every session. It needs tests. But the model is non-deterministic, slow and costs money, so you cannot test the harness by running the model. This lesson shows how to test everything around the model, with no model at all.
A scripted fake model
The key design choice is in Kite's model.py, which you write in the first build milestone: the loop depends on a Model with one method, complete(system, messages, tools). Anything with that method can stand in for the real model. Kite's ScriptedModel plays back a list of prepared replies and records what it was shown:
1class ScriptedModel:2 """Plays back prepared replies. Used by tests, and by replay in milestone 6."""34 def __init__(self, replies: list[Reply]):5 self.replies = list(replies)6 self.seen: list[list[dict]] = [] # the message list shown on each call78 def complete(self, system, messages, tools):9 self.seen.append(copy.deepcopy(messages))10 if not self.replies:11 raise RuntimeError("ScriptedModel has no replies left")12 return self.replies.pop(0)Two small helpers, use(name, **args) and say(text), build tool-call and text replies. With these, a test is a short script of what the model "does", followed by assertions about what the harness did in response.
The seen list is the important part. It lets a test check not only what the harness did, but what the model would have seen — which is the whole point of most harness features. A refusal is useless if the model never reads it.
What to test
Test the harness's decisions, one behaviour per test. Here are three from Kite's suite:
1def test_refusal_reaches_the_model_as_an_error(tmp_path):2 model = ScriptedModel([use("run", command="git push origin main"),3 say("Pushing needs a human. Stopping here.")])4 toolbox = Toolbox(tmp_path, Config(), check=Policy(refuse).check)5 outcome = run_session("Ship it.", model, toolbox, Config(), system="test")6 result = model.seen[1][-1]["content"][0] # last message on the 2nd call7 assert result["is_error"] and result["content"].startswith("Not permitted (deny)")8 assert outcome.status == "done"91011def test_cut_off_reply_is_not_done(tmp_path):12 model = ScriptedModel([Reply([{"type": "text", "text": "The fix is to"}], "max_tokens")])13 outcome = run_session("Fix it.", model, Toolbox(tmp_path, Config()), Config(), system="test")14 assert outcome.status == "cut_off"151617def test_gate_failures_are_capped(repo):18 (repo / "ledgerly" / "money.py").write_text("def add_fee(total, rate):\n return 0\n")19 model = ScriptedModel([say("Done.")] * 3)20 gate = Gate(repo, ["bash scripts/check.sh"], 60)21 outcome = run_session("Fix it.", model, Toolbox(repo, Config()), Config(), system="test", gate=gate)22 assert (outcome.status, outcome.turns) == ("gate_failed", 3)23 assert "exited with code 1" in model.seen[2][-1]["content"]The repo fixture creates a tiny stand-in for Ledgerly in a temporary folder: one module, one test, a check.sh, and a git history. You build it in milestone 4. The whole suite runs in under ten seconds and costs nothing.
A useful list of behaviours to cover:
| Area | Behaviours worth a test |
|---|---|
| Loop | Tool results go back with matching ids; every stop condition ends with the right status |
| Tools | Missing files, bad arguments and timeouts return errors instead of crashing; output is clipped |
| Permissions | Each tier with examples; chained commands are not "safe"; refusals explain themselves |
| Gate | Failing checks, missing tests, removed asserts; the attempt cap |
| Progress | Status changes only on verified outcomes; failed work is stashed, not lost |
The pycache bug would have been caught by a test like test_source_change_without_a_test_fails — but only if the fixture repository ran pytest as part of the check, as the real one does. Fixtures that are too clean hide bugs that real repositories expose. That is why Kite's fixture has a real check.sh that really runs pytest.
Replay: the real run, reproduced
Scripted tests cover behaviours you thought of. Replay covers the one that actually happened. Because Kite's event log stores every model reply, it can build a ScriptedModel from a log:
1def replay(path: Path) -> ScriptedModel:2 """A model that says, turn by turn, exactly what the recorded model said."""3 return ScriptedModel([Reply(e["content"], e["stop"], e["input_tokens"], e["output_tokens"])4 for e in read(path) if e["kind"] == "model"])To reproduce a failed session, check out the commit the session started from in a fresh worktree, and run Kite with --replay and the log file. The model's side of the conversation comes from the log. The harness's side — tool results, permission decisions, gate verdicts — runs live against the real files. You get the same run, deterministically, as many times as you like, for free.
This is most useful for testing a harness fix against the failure that prompted it. After fixing the pycache bug, the team replayed the session that had wrongly passed. With the fix, the replayed gate failed at the same point, with "no file under tests/", which is exactly the evidence they wanted.
Replay has a clear limit. It is exact only while the world matches the recording. If your harness change alters a tool result — say, a new permission rule refuses a command that the original run was allowed — the recorded model's next reply was written in response to the old result, and the replay stops making sense from that point. Use replay to check what changes up to the first divergence, not beyond it.
Golden runs
A golden run is a recorded session that you keep as a regression test. Pick three to five logs that cover important behaviour: a clean success, a gate failure that was fixed within the session, a permission refusal the agent worked around. Store each with the commit it started from. The test replays each one and asserts the outcome:
1def test_golden_late_fee_run(ledgerly_at_start, monkeypatch):2 progress.save(ledgerly_at_start / "kite-progress.json", F1_ONLY)3 log = Path("tests/golden/f1-late-fee.jsonl")4 assert main([str(ledgerly_at_start), "--unattended", "--replay", str(log)]) == 05 replayed = next((ledgerly_at_start / ".kite" / "runs").glob("*-F1-replay.jsonl"))6 assert [e["passed"] for e in read(replayed) if e["kind"] == "gate"] == [False, True]Here ledgerly_at_start is a fixture that checks out the recorded starting commit, and F1_ONLY is the progress data for that run. Kite names a replay's log with a -replay suffix, so the test can find it. When a harness change makes a golden test fail, it does not necessarily mean the change is wrong. It means the change alters behaviour on a real run, and a human should look at how before merging.
Tests are not evals
None of this tells you how good the agent is. Tests with fake models check that the harness does what you designed. How often the real model plus your harness finishes real tasks is a different question, answered by running the real model on a set of tasks and measuring the pass rate. That costs money and varies between runs, so do it on a schedule — nightly or weekly, on 10 to 20 representative tasks — not on every commit. Harness tests on every commit, evaluations on a schedule: each answers a question the other cannot.
Check your understanding
0 of 3 answered
1.Why do Kite's tests check model.seen and not only the outcome status?
2.You add a permission rule that denies pip install. You replay a recorded run in which the agent installed a package at turn 6. What should you expect?
3.What is the difference between the harness test suite and an evaluation run?