Harness Engineering: Making Coding Agents Dependable

Green tests that never touched your change


The late-fee session passed the gate. Lint clean, 214 tests green, app boots. A reviewer opened the diff and found the new fee code in ledgerly/invoices/service.py — 23 lines — and no new test. The suite was green because no test in it ever reached an overdue invoice. The 214 tests would have passed just the same if the fee code were deleted, or if it charged 200% instead of 2%.

A gate built from "the tests pass" answers the question "did anything that is tested break?". It does not answer "is the new behaviour correct?", and those are very different questions. The first lesson on gates made "done" a claim. This lesson makes sure the proof actually covers the claim.

The same late-fee session, two readingsThe green gate says• Lint clean• 214 of 214 tests pass• The app boots• Looks doneChanged-line coverage says• 2 of 23 new lines ran• Only the def lines, at import• No test reached an overdue invoice• The 30-versus-31 bug is hidden
Measuring whether tests ran the changed lines exposed the missing tests, and the new tests exposed an off-by-one at 30 days.

Three ways green lies

  • The tests never reach the change. The new code has no test, or the tests that exist take a path around it. Coverage of the change is zero, and the gate cannot tell.
  • The tests were weakened. The agent edited an assertion from assert total == Decimal("1180.00") to assert total > 0, or added a skip marker, or deleted a test that "was testing old behaviour".
  • The tests test the mock. The agent mocked apply_late_fee inside the test for apply_late_fee. Everything passes, and nothing real runs.

Each of these happens in real agent runs, and each gets more likely the longer the agent has been failing. A model under pressure to reach green, with tests in reach, will sometimes take the shortest path.

Rule 1: changed source needs a changed test

The cheapest check is structural: if the diff changes a source file, it must also add or change a test file. Kite's gate does exactly this:

Python
python = [f for f in changed if f.endswith(".py")]    # ignores .pyc and data filessource = sorted(f for f in python if not f.startswith("tests/"))tests = sorted(f for f in python if f.startswith("tests/"))if source and not tests:    problems.append(f"You changed {', '.join(source)} but no file under tests/. "                    "Add or update a test that fails without your change and passes with it.")

The first line has a story. An earlier version checked any changed path under tests/. It passed a session that wrote no tests at all, because running pytest had created new tests/__pycache__/*.pyc files, and those counted as "changed tests". The gate was fooled by its own side effect. This is the lesson of this whole page in miniature: a check that passes is only as good as what it actually looks at.

This rule is crude. It is satisfied by a test file that does not test the change, and it wrongly fails a pure refactor that should not need new tests. It is still worth having, because it costs nothing and catches the most common failure. For refactors, the task text can say so, and a human can accept the session.

Rule 2: the tests must run the changed lines

A stronger check asks: of the lines this diff added, how many did the test suite actually execute? coverage.py records which lines run; git diff says which lines are new. Put them together:

Python
def added_lines(base: str = "HEAD") -> dict[str, set[int]]:    subprocess.run(["git", "add", "--intent-to-add", "."], check=True)  # new files show in the diff    diff = subprocess.run(["git", "diff", "-U0", base, "--", "*.py", ":!tests/*"],                          capture_output=True, text=True, check=True).stdout    added, current = {}, None    for line in diff.splitlines():        if line.startswith("+++ b/"):            current = line[6:]        elif line.startswith("@@") and current:            start, count = re.search(r"\+(\d+)(?:,(\d+))?", line).groups()            first, n = int(start), int(count) if count is not None else 1            added.setdefault(current, set()).update(range(first, first + n))    return addeddef main(threshold: float = 0.8) -> int:    subprocess.run(["coverage", "run", "--source=ledgerly", "-m", "pytest", "-q", "-m", "not flaky"])    subprocess.run(["coverage", "json", "-q", "-o", "coverage.json"], check=True)    files = json.load(open("coverage.json"))["files"]    hit = total = 0    for path, lines in added_lines().items():        info = files.get(path, {"executed_lines": [], "missing_lines": []})        ran, missed = lines & set(info["executed_lines"]), lines & set(info["missing_lines"])        hit, total = hit + len(ran), total + len(ran) + len(missed)        for n in sorted(missed):            print(f"not run by any test: {path}:{n}")    print(f"changed lines run by tests: {hit}/{total}")    return 0 if total == 0 or hit / total >= threshold else 1

Save it as scripts/changed_coverage.py with import json, re, subprocess, sys at the top and sys.exit(main()) at the bottom. git diff -U0 prints each change with no context lines, and each @@ -a,b +c,d @@ header says that lines c to c+d-1 are new. Blank lines and comments are neither executed nor missing in coverage's view, so they do not count either way.

Two details were learned the hard way. The --intent-to-add line makes brand-new files show up in git diff; without it, a whole new module would be invisible. And --source=ledgerly makes coverage report files that no test ever imported. Without it, a new ledgerly/late_fees.py that no test touches simply does not appear in the report, and the script happily prints 0/0 and passes. On the unguarded late-fee session, this script prints changed lines run by tests: 2/23 — the two def lines, which run at import time — and fails.

The cost: coverage makes Ledgerly's suite about 30% slower, 38 seconds becomes 50. Run it in the gate only, not in the agent's inner loop.

Rule 3: nobody weakens the tests

The last rule protects the tests themselves. Kite's gate scans the diff of tests/ for removed lines that contain assert:

Python
removed = [line for line in git(self.root, "diff", "HEAD", "--", "tests/").splitlines()           if line.startswith("-") and not line.startswith("---") and "assert" in line]

If any exist, the gate fails and shows them. The agent may have a good reason — the task changed a behaviour, so an old assertion is now wrong — and then a human accepts it. But the default is that an agent does not delete evidence. You can extend the same scan to added lines containing pytest.mark.skip, pytest.skip( or xfail, and to new mock.patch targets that name the module under test.

Cheap structural rules

  • Changed source needs a changed test
  • No removed asserts, no new skips
  • Milliseconds; no extra tooling
  • Easy to satisfy without real testing

Execution evidence

  • Changed lines must be run by tests
  • Optionally: tests must fail when the change is reverted
  • Adds 10–30% to test time
  • Much harder to satisfy by accident

The strongest check of all is the revert test: hide the source change, keep the new tests, and run them. If they still pass, they do not test the change. It is slow and fiddly to automate, so most teams run it by hand in review, or on a sample of sessions:

Bash
git stash push -- ledgerly/                     # hide the source change; keep the new testspython -m pytest -q tests/unit/test_late_fee.py && echo "WARNING: tests pass without the change"git stash pop

Check your understanding

0 of 3 answered

1.The gate passes, but changed-line coverage for a session is 2/23, and the two covered lines are def lines. What does that tell you?

2.A changed-lines coverage script reports 0/0 and passes for a session that added a brand-new module ledgerly/late_fees.py. What is the most likely cause?

3.An agent's diff removes assert invoice.total == Decimal("1180.00") from a test. When is that acceptable?