Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Green tests that never touched your change
The late-fee session passed the gate. Lint clean, 214 tests green, app boots. A reviewer opened the diff and found the new fee code in ledgerly/invoices/service.py — 23 lines — and no new test. The suite was green because no test in it ever reached an overdue invoice. The 214 tests would have passed just the same if the fee code were deleted, or if it charged 200% instead of 2%.
A gate built from "the tests pass" answers the question "did anything that is tested break?". It does not answer "is the new behaviour correct?", and those are very different questions. The first lesson on gates made "done" a claim. This lesson makes sure the proof actually covers the claim.
Three ways green lies
- The tests never reach the change. The new code has no test, or the tests that exist take a path around it. Coverage of the change is zero, and the gate cannot tell.
- The tests were weakened. The agent edited an assertion from
assert total == Decimal("1180.00")toassert total > 0, or added a skip marker, or deleted a test that "was testing old behaviour". - The tests test the mock. The agent mocked
apply_late_feeinside the test forapply_late_fee. Everything passes, and nothing real runs.
Each of these happens in real agent runs, and each gets more likely the longer the agent has been failing. A model under pressure to reach green, with tests in reach, will sometimes take the shortest path.
Rule 1: changed source needs a changed test
The cheapest check is structural: if the diff changes a source file, it must also add or change a test file. Kite's gate does exactly this:
1python = [f for f in changed if f.endswith(".py")] # ignores .pyc and data files2source = sorted(f for f in python if not f.startswith("tests/"))3tests = sorted(f for f in python if f.startswith("tests/"))4if source and not tests:5 problems.append(f"You changed {', '.join(source)} but no file under tests/. "6 "Add or update a test that fails without your change and passes with it.")The first line has a story. An earlier version checked any changed path under tests/. It passed a session that wrote no tests at all, because running pytest had created new tests/__pycache__/*.pyc files, and those counted as "changed tests". The gate was fooled by its own side effect. This is the lesson of this whole page in miniature: a check that passes is only as good as what it actually looks at.
This rule is crude. It is satisfied by a test file that does not test the change, and it wrongly fails a pure refactor that should not need new tests. It is still worth having, because it costs nothing and catches the most common failure. For refactors, the task text can say so, and a human can accept the session.
Rule 2: the tests must run the changed lines
A stronger check asks: of the lines this diff added, how many did the test suite actually execute? coverage.py records which lines run; git diff says which lines are new. Put them together:
1def added_lines(base: str = "HEAD") -> dict[str, set[int]]:2 subprocess.run(["git", "add", "--intent-to-add", "."], check=True) # new files show in the diff3 diff = subprocess.run(["git", "diff", "-U0", base, "--", "*.py", ":!tests/*"],4 capture_output=True, text=True, check=True).stdout5 added, current = {}, None6 for line in diff.splitlines():7 if line.startswith("+++ b/"):8 current = line[6:]9 elif line.startswith("@@") and current:10 start, count = re.search(r"\+(\d+)(?:,(\d+))?", line).groups()11 first, n = int(start), int(count) if count is not None else 112 added.setdefault(current, set()).update(range(first, first + n))13 return added141516def main(threshold: float = 0.8) -> int:17 subprocess.run(["coverage", "run", "--source=ledgerly", "-m", "pytest", "-q", "-m", "not flaky"])18 subprocess.run(["coverage", "json", "-q", "-o", "coverage.json"], check=True)19 files = json.load(open("coverage.json"))["files"]20 hit = total = 021 for path, lines in added_lines().items():22 info = files.get(path, {"executed_lines": [], "missing_lines": []})23 ran, missed = lines & set(info["executed_lines"]), lines & set(info["missing_lines"])24 hit, total = hit + len(ran), total + len(ran) + len(missed)25 for n in sorted(missed):26 print(f"not run by any test: {path}:{n}")27 print(f"changed lines run by tests: {hit}/{total}")28 return 0 if total == 0 or hit / total >= threshold else 1Save it as scripts/changed_coverage.py with import json, re, subprocess, sys at the top and sys.exit(main()) at the bottom. git diff -U0 prints each change with no context lines, and each @@ -a,b +c,d @@ header says that lines c to c+d-1 are new. Blank lines and comments are neither executed nor missing in coverage's view, so they do not count either way.
Two details were learned the hard way. The --intent-to-add line makes brand-new files show up in git diff; without it, a whole new module would be invisible. And --source=ledgerly makes coverage report files that no test ever imported. Without it, a new ledgerly/late_fees.py that no test touches simply does not appear in the report, and the script happily prints 0/0 and passes. On the unguarded late-fee session, this script prints changed lines run by tests: 2/23 — the two def lines, which run at import time — and fails.
The cost: coverage makes Ledgerly's suite about 30% slower, 38 seconds becomes 50. Run it in the gate only, not in the agent's inner loop.
Rule 3: nobody weakens the tests
The last rule protects the tests themselves. Kite's gate scans the diff of tests/ for removed lines that contain assert:
removed = [line for line in git(self.root, "diff", "HEAD", "--", "tests/").splitlines() if line.startswith("-") and not line.startswith("---") and "assert" in line]If any exist, the gate fails and shows them. The agent may have a good reason — the task changed a behaviour, so an old assertion is now wrong — and then a human accepts it. But the default is that an agent does not delete evidence. You can extend the same scan to added lines containing pytest.mark.skip, pytest.skip( or xfail, and to new mock.patch targets that name the module under test.
Cheap structural rules
- Changed source needs a changed test
- No removed asserts, no new skips
- Milliseconds; no extra tooling
- Easy to satisfy without real testing
Execution evidence
- Changed lines must be run by tests
- Optionally: tests must fail when the change is reverted
- Adds 10–30% to test time
- Much harder to satisfy by accident
The strongest check of all is the revert test: hide the source change, keep the new tests, and run them. If they still pass, they do not test the change. It is slow and fiddly to automate, so most teams run it by hand in review, or on a sample of sessions:
git stash push -- ledgerly/ # hide the source change; keep the new testspython -m pytest -q tests/unit/test_late_fee.py && echo "WARNING: tests pass without the change"git stash popCheck your understanding
0 of 3 answered
1.The gate passes, but changed-line coverage for a session is 2/23, and the two covered lines are def lines. What does that tell you?
2.A changed-lines coverage script reports 0/0 and passes for a session that added a brand-new module ledgerly/late_fees.py. What is the most likely cause?
3.An agent's diff removes assert invoice.total == Decimal("1180.00") from a test. When is that acceptable?