Harness Engineering: Making Coding Agents Dependable

Milestone 4: the completion gate


Kite can now work on a real repository without doing damage. It can still declare victory falsely. The lesson on completion gates measured this on Ledgerly: 22% of unguarded "done" claims failed the full checks. This milestone makes the harness, not the model, decide when a session is finished.

You write kite/gate.py, make the one change to the loop that the whole build needs, and add a test fixture — a tiny Ledgerly with a real git history and a real check script — that the rest of the build's tests will use.

Two claims of done in one sessionpassFAILpasspasspasspasspasspasscheck.shTest changedAsserts keptverifyClaim at turn 3Claim at turn 6The failed row's feedback became the next user message.
With the gate in the loop, a reply with no tool calls is only a claim, and the evidence rules catch a change with no test.

The gate

Python
"""The completion gate: "done" is a claim, and these checks are the proof."""import subprocessfrom dataclasses import dataclassfrom pathlib import Pathfrom .tools import clip@dataclassclass Verdict:    passed: bool    feedback: strdef git(root: Path, *args: str) -> str:    return subprocess.run(["git", *args], cwd=root, capture_output=True, text=True).stdoutclass Gate:    def __init__(self, root: Path, checks: list[str], timeout: int):        self.root, self.checks, self.timeout = root, checks, timeout    def check(self) -> Verdict:        problems = [p for cmd in self.checks if (p := self.run_check(cmd))]        problems += self.evidence_problems()        if not problems:            return Verdict(True, "All completion checks passed.")        return Verdict(False, "You said the task is done, but the completion gate failed.\n\n"                       + "\n\n".join(problems)                       + "\n\nFix the cause, then finish again. Do not delete, skip or weaken tests.")    def run_check(self, cmd: str) -> str | None:        try:            proc = subprocess.run(cmd, shell=True, cwd=self.root, capture_output=True,                                  text=True, timeout=self.timeout)        except subprocess.TimeoutExpired:            return f"'{cmd}' timed out after {self.timeout}s."        if proc.returncode == 0:            return None        return f"'{cmd}' exited with code {proc.returncode}:\n{clip(proc.stdout + proc.stderr, 3000)}"

The gate runs every check command and collects the problems. It runs all of them rather than stopping at the first: the agent is about to spend a turn fixing things, so it may as well learn about every failure at once. Put fast checks first in check.sh itself, where set -e stops early. Each failure is clipped to 3,000 characters, and the whole message follows the shape from the gate lesson: what failed, the evidence, and what not to do.

Now the evidence rules from the lesson on what a passing run proves:

Python
    def evidence_problems(self) -> list[str]:        changed = set(git(self.root, "diff", "--name-only", "HEAD").split())        changed |= set(git(self.root, "ls-files", "--others", "--exclude-standard").split())        python = [f for f in changed if f.endswith(".py")]    # ignores .pyc and data files        source = sorted(f for f in python if not f.startswith("tests/"))        tests = sorted(f for f in python if f.startswith("tests/"))        problems = []        if source and not tests:            problems.append(f"You changed {', '.join(source)} but no file under tests/. "                            "Add or update a test that fails without your change and passes with it.")        removed = [line for line in git(self.root, "diff", "HEAD", "--", "tests/").splitlines()                   if line.startswith("-") and not line.startswith("---") and "assert" in line]        if removed:            problems.append(f"You removed {len(removed)} assert line(s) from tests:\n"                            + "\n".join(removed[:5]))        return problems

"Changed" means tracked files that differ from the last commit, plus new files that are not ignored. This is why every session must start from a clean tree, which milestone 5 guarantees: then everything in the diff belongs to this session. The .py filter is the fix for the __pycache__ bug you met in the testing lesson.

The one change to the loop

Replace run_session in kite/loop.py with this version. It is the final form of the loop; milestone 6 adds logging without touching it.

Python
def run_session(task: str, model, toolbox, cfg, *, system: str, gate=None) -> Outcome:    messages = [{"role": "user", "content": task}]    tokens = gate_failures = 0    for turn in range(1, cfg.max_turns + 1):        reply = model.complete(system, messages, toolbox.specs())        tokens += reply.input_tokens + reply.output_tokens        messages.append({"role": "assistant", "content": reply.content})        if reply.stop_reason in ("max_tokens", "refusal"):            return Outcome("cut_off", turn, tokens, reply.stop_reason)        calls = reply.tool_calls        if calls:            results = [toolbox.run(call) for call in calls]            messages.append({"role": "user", "content": [r.to_block() for r in results]})        elif gate is None:            return Outcome("done", turn, tokens, reply.text)        else:            verdict = gate.check()            if verdict.passed:                return Outcome("done", turn, tokens, reply.text)            gate_failures += 1            if gate_failures >= cfg.max_gate_failures:                return Outcome("gate_failed", turn, tokens, verdict.feedback)            messages.append({"role": "user", "content": verdict.feedback})        if tokens >= cfg.token_budget:            return Outcome("out_of_tokens", turn, tokens, reply.text)    return Outcome("out_of_turns", cfg.max_turns, tokens, "")

When the model stops calling tools, the loop no longer returns "done" straight away. If there is a gate, it runs it. A pass ends the session as done. A failure goes back to the model as the next user message — the message list stays valid, because an assistant turn is followed by a user turn — and the loop continues. After three failures, the session ends as gate_failed, and the summary is the gate's feedback, so the next session learns exactly what blocked this one.

The gate is optional (gate=None), so the tests from milestone 1 still pass unchanged. In __main__.py, add from .gate import Gate, create gate = Gate(root, cfg.checks, cfg.command_timeout) after the toolbox, and pass gate=gate to run_session.

What the gate costs, and what it cannot see

On Ledgerly, one gate run takes about 45 seconds: 2 for lint, 38 for the tests, 3 for the app boot check, a moment for git. With at most three refusals and one pass, the gate adds at most three minutes to a session that usually runs ten to twenty. The timeout on each check comes from cfg.command_timeout, 180 seconds, so a hung test cannot hang the harness.

Be clear about what the gate does not check. It does not know whether the change does what the task meant: a late fee of 2% per day would pass, if the agent wrote a test that expects it. It does not check scope beyond the test rules. And its evidence rules are cheap heuristics, not proof of coverage. The gate removes a whole class of false "done" claims. It does not remove the need for a human to read the diff before it is merged. It makes that review shorter, because the reviewer no longer has to check whether the tests pass.

A tiny Ledgerly for tests

The gate runs git and shell commands, so its tests need a real git repository. Create tests/conftest.py:

Python
import subprocessimport pytestMONEY = "def add_fee(total, rate):\n    return round(total * rate / 100, 2)\n"TEST = ("from ledgerly.money import add_fee\n\n\ndef test_fee():\n"        "    assert add_fee(1000, 2) == 20\n")@pytest.fixturedef repo(tmp_path):    """A tiny stand-in for Ledgerly: one module, one test, one check script, git history."""    (tmp_path / "ledgerly").mkdir()    (tmp_path / "ledgerly" / "__init__.py").write_text("")    (tmp_path / "ledgerly" / "money.py").write_text(MONEY)    (tmp_path / "tests").mkdir()    (tmp_path / "tests" / "test_money.py").write_text(TEST)    (tmp_path / ".gitignore").write_text("__pycache__/\n")    (tmp_path / "scripts").mkdir()    (tmp_path / "scripts" / "check.sh").write_text("set -e\npython -m pytest -q\n")    for cmd in (["git", "init", "-q"], ["git", "add", "-A"],                ["git", "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", "init"]):        subprocess.run(cmd, cwd=tmp_path, check=True)    subprocess.run(["git", "config", "user.name", "Kite"], cwd=tmp_path)    subprocess.run(["git", "config", "user.email", "kite@example.com"], cwd=tmp_path)    return tmp_path

It is deliberately realistic in the ways that matter. check.sh really runs pytest, so running the gate creates __pycache__ folders just as it does on Ledgerly — the fixture would have caught the bug from the testing lesson. It has a .gitignore, a first commit, and a git identity for the commits that milestone 5 makes.

Now tests/test_gate.py:

Python
from kite.gate import Gatedef test_passes_on_a_clean_tree(repo):    assert Gate(repo, ["bash scripts/check.sh"], 60).check().passeddef test_source_change_without_a_test_fails(repo):    (repo / "ledgerly" / "money.py").write_text("def add_fee(total, rate):\n    return total * rate / 100\n")    verdict = Gate(repo, ["bash scripts/check.sh"], 60).check()    assert not verdict.passed and "no file under tests/" in verdict.feedbackdef test_failing_check_and_removed_assert_are_reported(repo):    (repo / "tests" / "test_money.py").write_text("def test_fee():\n    assert 1 == 2\n")    verdict = Gate(repo, ["bash scripts/check.sh"], 60).check()    assert "exited with code 1" in verdict.feedback    assert "removed 1 assert line" in verdict.feedback

The third test is worth reading closely. The "agent" replaced a real assertion with assert 1 == 2. The gate reports both the failing check and the removed assertion — two separate problems, both visible to the model at once.

Check your understanding

0 of 3 answered

1.Why does Gate.check run every check command instead of stopping at the first failure?

2.After the loop change, what does the model receive when the gate fails?

3.Why does the fixture's check.sh run pytest for real instead of just exiting 0?