Harness Engineering: Making Coding Agents Dependable

Testing the harness itself


The Ledgerly team once tightened Kite's gate with a new rule: a changed source file must come with a changed test file. The change looked obviously right, and it was merged in an afternoon. That night, a session that wrote no test at all passed the gate. Running pytest had created new tests/__pycache__/*.pyc files, and the rule counted them as "changed tests". A check designed to catch missing tests was fooled by the side effect of running them.

The harness is code. It has bugs like any code, and its bugs are expensive, because a wrong gate or a wrong permission rule quietly affects every session. It needs tests. But the model is non-deterministic, slow and costs money, so you cannot test the harness by running the model. This lesson shows how to test everything around the model, with no model at all.

Two ways to test without the real modelScripted fake model• Tests behaviours you designed• No key, no tokens, about 8 seconds• Checks what the model would see• Runs on every commitReplay from a log• Reproduces the run that happened• Model side recorded, harness side live• Proves a fix against a real failure• Exact only until the first divergence
A gate rule fooled by pycache files passed review; a fixture that really runs pytest would have caught it on the first run.

A scripted fake model

The key design choice is in Kite's model.py, which you write in the first build milestone: the loop depends on a Model with one method, complete(system, messages, tools). Anything with that method can stand in for the real model. Kite's ScriptedModel plays back a list of prepared replies and records what it was shown:

Python
class ScriptedModel:    """Plays back prepared replies. Used by tests, and by replay in milestone 6."""    def __init__(self, replies: list[Reply]):        self.replies = list(replies)        self.seen: list[list[dict]] = []     # the message list shown on each call    def complete(self, system, messages, tools):        self.seen.append(copy.deepcopy(messages))        if not self.replies:            raise RuntimeError("ScriptedModel has no replies left")        return self.replies.pop(0)

Two small helpers, use(name, **args) and say(text), build tool-call and text replies. With these, a test is a short script of what the model "does", followed by assertions about what the harness did in response.

The seen list is the important part. It lets a test check not only what the harness did, but what the model would have seen — which is the whole point of most harness features. A refusal is useless if the model never reads it.

What to test

Test the harness's decisions, one behaviour per test. Here are three from Kite's suite:

Python
def test_refusal_reaches_the_model_as_an_error(tmp_path):    model = ScriptedModel([use("run", command="git push origin main"),                           say("Pushing needs a human. Stopping here.")])    toolbox = Toolbox(tmp_path, Config(), check=Policy(refuse).check)    outcome = run_session("Ship it.", model, toolbox, Config(), system="test")    result = model.seen[1][-1]["content"][0]            # last message on the 2nd call    assert result["is_error"] and result["content"].startswith("Not permitted (deny)")    assert outcome.status == "done"def test_cut_off_reply_is_not_done(tmp_path):    model = ScriptedModel([Reply([{"type": "text", "text": "The fix is to"}], "max_tokens")])    outcome = run_session("Fix it.", model, Toolbox(tmp_path, Config()), Config(), system="test")    assert outcome.status == "cut_off"def test_gate_failures_are_capped(repo):    (repo / "ledgerly" / "money.py").write_text("def add_fee(total, rate):\n    return 0\n")    model = ScriptedModel([say("Done.")] * 3)    gate = Gate(repo, ["bash scripts/check.sh"], 60)    outcome = run_session("Fix it.", model, Toolbox(repo, Config()), Config(), system="test", gate=gate)    assert (outcome.status, outcome.turns) == ("gate_failed", 3)    assert "exited with code 1" in model.seen[2][-1]["content"]

The repo fixture creates a tiny stand-in for Ledgerly in a temporary folder: one module, one test, a check.sh, and a git history. You build it in milestone 4. The whole suite runs in under ten seconds and costs nothing.

A useful list of behaviours to cover:

AreaBehaviours worth a test
LoopTool results go back with matching ids; every stop condition ends with the right status
ToolsMissing files, bad arguments and timeouts return errors instead of crashing; output is clipped
PermissionsEach tier with examples; chained commands are not "safe"; refusals explain themselves
GateFailing checks, missing tests, removed asserts; the attempt cap
ProgressStatus changes only on verified outcomes; failed work is stashed, not lost

The pycache bug would have been caught by a test like test_source_change_without_a_test_fails — but only if the fixture repository ran pytest as part of the check, as the real one does. Fixtures that are too clean hide bugs that real repositories expose. That is why Kite's fixture has a real check.sh that really runs pytest.

Replay: the real run, reproduced

Scripted tests cover behaviours you thought of. Replay covers the one that actually happened. Because Kite's event log stores every model reply, it can build a ScriptedModel from a log:

Python
def replay(path: Path) -> ScriptedModel:    """A model that says, turn by turn, exactly what the recorded model said."""    return ScriptedModel([Reply(e["content"], e["stop"], e["input_tokens"], e["output_tokens"])                          for e in read(path) if e["kind"] == "model"])

To reproduce a failed session, check out the commit the session started from in a fresh worktree, and run Kite with --replay and the log file. The model's side of the conversation comes from the log. The harness's side — tool results, permission decisions, gate verdicts — runs live against the real files. You get the same run, deterministically, as many times as you like, for free.

This is most useful for testing a harness fix against the failure that prompted it. After fixing the pycache bug, the team replayed the session that had wrongly passed. With the fix, the replayed gate failed at the same point, with "no file under tests/", which is exactly the evidence they wanted.

Replay has a clear limit. It is exact only while the world matches the recording. If your harness change alters a tool result — say, a new permission rule refuses a command that the original run was allowed — the recorded model's next reply was written in response to the old result, and the replay stops making sense from that point. Use replay to check what changes up to the first divergence, not beyond it.

Golden runs

A golden run is a recorded session that you keep as a regression test. Pick three to five logs that cover important behaviour: a clean success, a gate failure that was fixed within the session, a permission refusal the agent worked around. Store each with the commit it started from. The test replays each one and asserts the outcome:

Python
def test_golden_late_fee_run(ledgerly_at_start, monkeypatch):    progress.save(ledgerly_at_start / "kite-progress.json", F1_ONLY)    log = Path("tests/golden/f1-late-fee.jsonl")    assert main([str(ledgerly_at_start), "--unattended", "--replay", str(log)]) == 0    replayed = next((ledgerly_at_start / ".kite" / "runs").glob("*-F1-replay.jsonl"))    assert [e["passed"] for e in read(replayed) if e["kind"] == "gate"] == [False, True]

Here ledgerly_at_start is a fixture that checks out the recorded starting commit, and F1_ONLY is the progress data for that run. Kite names a replay's log with a -replay suffix, so the test can find it. When a harness change makes a golden test fail, it does not necessarily mean the change is wrong. It means the change alters behaviour on a real run, and a human should look at how before merging.

Tests are not evals

None of this tells you how good the agent is. Tests with fake models check that the harness does what you designed. How often the real model plus your harness finishes real tasks is a different question, answered by running the real model on a set of tasks and measuring the pass rate. That costs money and varies between runs, so do it on a schedule — nightly or weekly, on 10 to 20 representative tasks — not on every commit. Harness tests on every commit, evaluations on a schedule: each answers a question the other cannot.

Check your understanding

0 of 3 answered

1.Why do Kite's tests check model.seen and not only the outcome status?

2.You add a permission rule that denies pip install. You replay a recorded run in which the agent installed a package at turn 6. What should you expect?

3.What is the difference between the harness test suite and an evaluation run?