Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Milestone 4: the completion gate
Kite can now work on a real repository without doing damage. It can still declare victory falsely. The lesson on completion gates measured this on Ledgerly: 22% of unguarded "done" claims failed the full checks. This milestone makes the harness, not the model, decide when a session is finished.
You write kite/gate.py, make the one change to the loop that the whole build needs, and add a test fixture — a tiny Ledgerly with a real git history and a real check script — that the rest of the build's tests will use.
The gate
1"""The completion gate: "done" is a claim, and these checks are the proof."""2import subprocess3from dataclasses import dataclass4from pathlib import Path56from .tools import clip789@dataclass10class Verdict:11 passed: bool12 feedback: str131415def git(root: Path, *args: str) -> str:16 return subprocess.run(["git", *args], cwd=root, capture_output=True, text=True).stdout171819class Gate:20 def __init__(self, root: Path, checks: list[str], timeout: int):21 self.root, self.checks, self.timeout = root, checks, timeout2223 def check(self) -> Verdict:24 problems = [p for cmd in self.checks if (p := self.run_check(cmd))]25 problems += self.evidence_problems()26 if not problems:27 return Verdict(True, "All completion checks passed.")28 return Verdict(False, "You said the task is done, but the completion gate failed.\n\n"29 + "\n\n".join(problems)30 + "\n\nFix the cause, then finish again. Do not delete, skip or weaken tests.")3132 def run_check(self, cmd: str) -> str | None:33 try:34 proc = subprocess.run(cmd, shell=True, cwd=self.root, capture_output=True,35 text=True, timeout=self.timeout)36 except subprocess.TimeoutExpired:37 return f"'{cmd}' timed out after {self.timeout}s."38 if proc.returncode == 0:39 return None40 return f"'{cmd}' exited with code {proc.returncode}:\n{clip(proc.stdout + proc.stderr, 3000)}"The gate runs every check command and collects the problems. It runs all of them rather than stopping at the first: the agent is about to spend a turn fixing things, so it may as well learn about every failure at once. Put fast checks first in check.sh itself, where set -e stops early. Each failure is clipped to 3,000 characters, and the whole message follows the shape from the gate lesson: what failed, the evidence, and what not to do.
Now the evidence rules from the lesson on what a passing run proves:
1 def evidence_problems(self) -> list[str]:2 changed = set(git(self.root, "diff", "--name-only", "HEAD").split())3 changed |= set(git(self.root, "ls-files", "--others", "--exclude-standard").split())4 python = [f for f in changed if f.endswith(".py")] # ignores .pyc and data files5 source = sorted(f for f in python if not f.startswith("tests/"))6 tests = sorted(f for f in python if f.startswith("tests/"))7 problems = []8 if source and not tests:9 problems.append(f"You changed {', '.join(source)} but no file under tests/. "10 "Add or update a test that fails without your change and passes with it.")11 removed = [line for line in git(self.root, "diff", "HEAD", "--", "tests/").splitlines()12 if line.startswith("-") and not line.startswith("---") and "assert" in line]13 if removed:14 problems.append(f"You removed {len(removed)} assert line(s) from tests:\n"15 + "\n".join(removed[:5]))16 return problems"Changed" means tracked files that differ from the last commit, plus new files that are not ignored. This is why every session must start from a clean tree, which milestone 5 guarantees: then everything in the diff belongs to this session. The .py filter is the fix for the __pycache__ bug you met in the testing lesson.
The one change to the loop
Replace run_session in kite/loop.py with this version. It is the final form of the loop; milestone 6 adds logging without touching it.
1def run_session(task: str, model, toolbox, cfg, *, system: str, gate=None) -> Outcome:2 messages = [{"role": "user", "content": task}]3 tokens = gate_failures = 04 for turn in range(1, cfg.max_turns + 1):5 reply = model.complete(system, messages, toolbox.specs())6 tokens += reply.input_tokens + reply.output_tokens7 messages.append({"role": "assistant", "content": reply.content})8 if reply.stop_reason in ("max_tokens", "refusal"):9 return Outcome("cut_off", turn, tokens, reply.stop_reason)10 calls = reply.tool_calls11 if calls:12 results = [toolbox.run(call) for call in calls]13 messages.append({"role": "user", "content": [r.to_block() for r in results]})14 elif gate is None:15 return Outcome("done", turn, tokens, reply.text)16 else:17 verdict = gate.check()18 if verdict.passed:19 return Outcome("done", turn, tokens, reply.text)20 gate_failures += 121 if gate_failures >= cfg.max_gate_failures:22 return Outcome("gate_failed", turn, tokens, verdict.feedback)23 messages.append({"role": "user", "content": verdict.feedback})24 if tokens >= cfg.token_budget:25 return Outcome("out_of_tokens", turn, tokens, reply.text)26 return Outcome("out_of_turns", cfg.max_turns, tokens, "")When the model stops calling tools, the loop no longer returns "done" straight away. If there is a gate, it runs it. A pass ends the session as done. A failure goes back to the model as the next user message — the message list stays valid, because an assistant turn is followed by a user turn — and the loop continues. After three failures, the session ends as gate_failed, and the summary is the gate's feedback, so the next session learns exactly what blocked this one.
The gate is optional (gate=None), so the tests from milestone 1 still pass unchanged. In __main__.py, add from .gate import Gate, create gate = Gate(root, cfg.checks, cfg.command_timeout) after the toolbox, and pass gate=gate to run_session.
What the gate costs, and what it cannot see
On Ledgerly, one gate run takes about 45 seconds: 2 for lint, 38 for the tests, 3 for the app boot check, a moment for git. With at most three refusals and one pass, the gate adds at most three minutes to a session that usually runs ten to twenty. The timeout on each check comes from cfg.command_timeout, 180 seconds, so a hung test cannot hang the harness.
Be clear about what the gate does not check. It does not know whether the change does what the task meant: a late fee of 2% per day would pass, if the agent wrote a test that expects it. It does not check scope beyond the test rules. And its evidence rules are cheap heuristics, not proof of coverage. The gate removes a whole class of false "done" claims. It does not remove the need for a human to read the diff before it is merged. It makes that review shorter, because the reviewer no longer has to check whether the tests pass.
A tiny Ledgerly for tests
The gate runs git and shell commands, so its tests need a real git repository. Create tests/conftest.py:
1import subprocess23import pytest45MONEY = "def add_fee(total, rate):\n return round(total * rate / 100, 2)\n"6TEST = ("from ledgerly.money import add_fee\n\n\ndef test_fee():\n"7 " assert add_fee(1000, 2) == 20\n")8910@pytest.fixture11def repo(tmp_path):12 """A tiny stand-in for Ledgerly: one module, one test, one check script, git history."""13 (tmp_path / "ledgerly").mkdir()14 (tmp_path / "ledgerly" / "__init__.py").write_text("")15 (tmp_path / "ledgerly" / "money.py").write_text(MONEY)16 (tmp_path / "tests").mkdir()17 (tmp_path / "tests" / "test_money.py").write_text(TEST)18 (tmp_path / ".gitignore").write_text("__pycache__/\n")19 (tmp_path / "scripts").mkdir()20 (tmp_path / "scripts" / "check.sh").write_text("set -e\npython -m pytest -q\n")21 for cmd in (["git", "init", "-q"], ["git", "add", "-A"],22 ["git", "-c", "user.name=t", "-c", "user.email=t@t", "commit", "-qm", "init"]):23 subprocess.run(cmd, cwd=tmp_path, check=True)24 subprocess.run(["git", "config", "user.name", "Kite"], cwd=tmp_path)25 subprocess.run(["git", "config", "user.email", "kite@example.com"], cwd=tmp_path)26 return tmp_pathIt is deliberately realistic in the ways that matter. check.sh really runs pytest, so running the gate creates __pycache__ folders just as it does on Ledgerly — the fixture would have caught the bug from the testing lesson. It has a .gitignore, a first commit, and a git identity for the commits that milestone 5 makes.
Now tests/test_gate.py:
1from kite.gate import Gate234def test_passes_on_a_clean_tree(repo):5 assert Gate(repo, ["bash scripts/check.sh"], 60).check().passed678def test_source_change_without_a_test_fails(repo):9 (repo / "ledgerly" / "money.py").write_text("def add_fee(total, rate):\n return total * rate / 100\n")10 verdict = Gate(repo, ["bash scripts/check.sh"], 60).check()11 assert not verdict.passed and "no file under tests/" in verdict.feedback121314def test_failing_check_and_removed_assert_are_reported(repo):15 (repo / "tests" / "test_money.py").write_text("def test_fee():\n assert 1 == 2\n")16 verdict = Gate(repo, ["bash scripts/check.sh"], 60).check()17 assert "exited with code 1" in verdict.feedback18 assert "removed 1 assert line" in verdict.feedbackThe third test is worth reading closely. The "agent" replaced a real assertion with assert 1 == 2. The gate reports both the failing check and the removed assertion — two separate problems, both visible to the model at once.
Check your understanding
0 of 3 answered
1.Why does Gate.check run every check command instead of stopping at the first failure?
2.After the loop change, what does the model receive when the gate fails?
3.Why does the fixture's check.sh run pytest for real instead of just exiting 0?