Harness Engineering: Making Coding Agents Dependable

Done is a claim, verification is the proof


Over 50 late-fee, export and reminder tasks on Ledgerly, an unguarded agent said "done" 50 times. In 11 of those runs the full check failed. That is a 22% false-done rate, and every one of those runs ended with a confident, well-written summary.

The model was not lying. In each case it believed it had finished. It had run the tests it thought were relevant, and they passed. It simply did not run the ones that failed: a lint error in a file it had edited, a test in another module that used a function it changed, an app that no longer started because of a circular import. From inside the session, "I ran the tests" and "the tests pass" feel like the same thing.

The fix is structural: the harness, not the model, decides when a session is finished. The model's "done" becomes a request to be checked, and the check is something you define.

What happens when the agent says doneAgentclaims doneruff check — 2 sTests exceptflaky — 38 sApp boots — 3 sFeature's ownverify testA failure goes back as the next message; three failures end the session.
The gate took Ledgerly from a 22% false-done rate to zero accepted false claims, for 1.4 extra turns on average.

What goes into the gate

A completion gate is a list of commands that must all succeed before the harness accepts "done". For Ledgerly, the list lives in one script so that humans, CI and the agent all run the same thing:

Bash
#!/usr/bin/env bash# scripts/check.sh: the completion gate for agents, humans and CIset -euo pipefailruff check ledgerly tests                      # 2 s: style and obvious errorspython -m pytest -q -x -m "not flaky"          # 38 s: every test except known-flaky onesflask --app ledgerly.app routes > /dev/null    # 3 s: the app still imports and registers routes

Order matters. Put the fastest checks first, and stop at the first failure (set -e and pytest's -x). If lint fails in 2 seconds, there is no reason to spend 38 seconds on tests before telling the agent. The last line is a cheap "does the app start" check. It catches circular imports and broken configuration that no unit test notices.

On top of the shared script, each feature can add its own acceptance command. For the late fee, it is python -m pytest -q tests/unit/test_late_fee.py. This makes sure the specific behaviour was tested, not just that nothing else broke.

CheckTimeCatches
ruff check2 sUnused imports, undefined names, syntax errors
Test suite without flaky tests38 sBroken behaviour anywhere, including in other modules
App boots3 sImport cycles, broken configuration
Feature acceptance test1–5 sThe requested behaviour exists and is tested

About 45 seconds in total. A Ledgerly session makes 20 to 30 model calls at 5 to 15 seconds each, so the gate adds under 10% to the session time. That is cheap for turning 22% false-done into zero accepted false-done.

The gate's message is a prompt

When the gate fails, its output goes back to the agent as the next user message. That makes the failure text one of the most important prompts in the harness. A good message says what failed, shows the evidence, and says what not to do:

Text
You said the task is done, but the completion gate failed.'bash scripts/check.sh' exited with code 1:ledgerly/invoices/service.py:14:1: F401 [*] `datetime.timezone` imported but unusedFound 1 error.Fix the cause, then finish again. Do not delete, skip or weaken tests.

Three details matter. First, the output is clipped: the start and, mostly, the end, where test summaries and tracebacks live. A 38,000-line failure dump buries the one line that matters. Second, the last sentence closes the most tempting shortcut. Without it, some fraction of agents will "fix" a failing test by editing its assertion. Third, the message is the same every time, so you can see in logs how agents respond to it and improve it.

Limit the number of attempts. Kite allows three failed gates per session, then stops with status gate_failed and keeps the work aside for a human. An agent that fails the same check three times is usually stuck on something it cannot see, and a fourth attempt rarely helps.

Flaky tests and the gate

A flaky test is poison for a gate. Ledgerly's PDF export test fails 1 run in 8 for reasons unrelated to any change: its headless browser renderer sometimes takes 6 seconds to start and hits a 5-second timeout. Inside a gate, that means one in eight good sessions fails for nothing. Worse, it teaches the agent that the way to pass is to make that test go away.

Handle flaky tests before they reach the gate, and never let the agent make the decision:

  1. Mark it — @pytest.mark.flaky on the test, with a comment saying why and a link to the ticket.
  2. Exclude it from the gate — -m "not flaky" in check.sh, so it cannot block anyone's work.
  3. Run it somewhere else — a nightly CI job runs flaky tests with retries and reports trends.
  4. Protect it — the instruction file says not to touch it, and the gate fails any diff that removes test assertions.
  5. Fix it for real — the quarantine is temporary; the ticket has an owner and a date.

The difference from what the unguarded agent did is who decides. The agent skipping a test to get green hides a signal. The team quarantining it, with a ticket and a separate run, keeps the signal and moves it somewhere it cannot block unrelated work.

Gates in tools you already use

You do not need Kite to have a gate. In Claude Code, a Stop hook can run your check script when the agent tries to finish. If the hook exits with code 2, the stop is blocked and the hook's error output is shown to the agent, which then keeps working. A minimal configuration in .claude/settings.json looks like this:

JSON
{  "hooks": {    "Stop": [      {"hooks": [{"type": "command", "command": "bash scripts/agent-gate.sh"}]}    ]  }}

Your agent-gate.sh runs check.sh, prints the clipped failure to standard error, and exits with 2 on failure. Read the JSON the hook receives on standard input: its stop_hook_active field is true when the agent is already continuing because of a stop hook. Count those rounds in a small state file and let the stop through after a few, instead of looping forever.

Check your understanding

0 of 3 answered

1.Why does Ledgerly's gate run ruff check before the test suite?

2.The gate keeps failing because the flaky PDF test fails 1 run in 8. What is the right response?

3.Why does Kite stop a session after three failed gates instead of letting the agent keep trying?