Harness Engineering: Making Coding Agents Dependable

Git as save points


After two weeks of agent sessions on Ledgerly, CI turned red. A test named test_interstate_split in tests/unit/test_gst.py was failing: invoices between two states were splitting GST wrongly by one paisa. Fourteen agent sessions had run since the last known good build. Which one did it?

If those fourteen sessions had produced one big, messy branch of work-in-progress commits, answering that question would take a morning of reading diffs. Because each session produced exactly one commit — one verified feature, or one recorded failure — the answer took four test runs and about three minutes.

Git is the one memory system every repository already has. This lesson uses it deliberately: as a series of save points that the harness creates, so you can find, undo or recover any session's work.

Bisecting fourteen one-feature commitsF1F2F3F4F5F6F7F8F9012345678goodfirst badbadgit bisect run needs about four test runs for fourteen commits.
One verified feature per commit turned a one-paisa GST regression into one named feature, F6, with its own diff and log.

One session, one commit

Kite commits at the end of every session, whatever the outcome. The rules are simple:

  • Done: commit everything — the feature's code, its tests, and the updated progress file — as kite: F2 done.
  • Not done: move the unfinished work into a git stash, then commit only the progress file update, as kite: F2 out_of_turns.

Here is the code, from Kite's progress.py:

Python
def finish(root: Path, path: Path, data: dict, feature: dict, outcome) -> str:    """Save the session's result and leave a clean tree; return the new commit hash."""    if outcome.status != "done":                       # keep failed work, but out of the way        subprocess.run(["git", "stash", "push", "-u", "-q", "-m", f"kite {feature['id']} failed"], cwd=root)    record(data, feature, outcome)    save(path, data)    subprocess.run(["git", "add", "-A"], cwd=root, check=True)    subprocess.run(["git", "commit", "-q", "-m", f"kite: {feature['id']} {outcome.status}"], cwd=root, check=True)    return subprocess.run(["git", "rev-parse", "--short", "HEAD"], cwd=root,                          capture_output=True, text=True).stdout.strip()

The order matters. On failure, the stash happens before the progress file is saved, so the stash holds exactly the agent's unfinished work and the commit holds exactly the harness's record. The next session starts from a clean tree, which also keeps the gate's "what changed" checks honest: every changed file belongs to the current session.

Failed work is not thrown away. It sits in the stash list with a clear label:

Bash
$ git stash liststash@{0}: On agent/late-fees: kite F2 failed$ git stash show -p stash@{0}       # read what the failed session did$ git stash apply stash@{0}         # bring it back, if a human wants to finish it

What a save point should record

A commit is only a useful save point if you can tell, months later, what it was and how it was checked. Kite's messages are short on purpose — kite: F6 done — because the rest of the story is one step away. The progress file in the same commit holds the feature's title, its verify command and the agent's summary. The event log for the session sits under .kite/runs/, named with the time and the feature id.

If your team reads history mostly through git log, extend the commit message with a body: the agent's summary, the number of turns and tokens, and the log file name. That costs one extra argument to git commit and saves a search later. Do not let the agent write commit messages freely, though. A message written by the model describes what it meant to do; a message written by the harness describes what was verified.

Where the commits go

Agents should never commit to your main branch. The Ledgerly setup gives each batch of work its own branch in its own worktree, as in the lesson on hard limits:

Bash
git worktree add ../ledgerly-agent -b agent/late-fees

Kite commits on agent/late-fees. A human reviews the branch and opens a pull request. Kite's permission layer denies git push, so nothing leaves the machine without a person deciding it should.

When the pull request is merged, you choose between keeping the per-session commits or squashing them into one. Keeping them makes git log longer and preserves the ability to bisect to a single feature. Squashing makes history tidier and loses that ability for this batch. Ledgerly keeps them; each commit is a whole, tested feature, so they read like a changelog anyway.

Rolling back safely

When a verified feature turns out to be wrong later, undo it with a new commit, not by rewriting history:

Bash
git revert --no-edit a41c0de        # a new commit that undoes F6, history intact

git revert is safe on a shared branch, and it keeps the record of what happened. git reset --hard rewrites history and throws away work, which is exactly why Kite's permission layer denies it to the agent. A human may still use it on an agent-only branch when they are sure, but it should be a rare, deliberate act.

Bisecting agent mistakes

git bisect finds the commit that broke something by binary search: you give it one bad commit and one good commit, and it checks out the commit in the middle, and so on. With 14 commits it needs about 4 steps, since each step halves the range. git bisect run automates it with any command that exits 0 for good and 1 for bad:

Bash
git bisect start HEAD 3f2a91c       # bad now; good at 3f2a91c, before the batch startedgit bisect run python -m pytest -q tests/unit/test_gst.py::test_interstate_splitgit bisect reset                    # go back to where you started

On Ledgerly, the output ended with a41c0de is the first bad commit and the message kite: F6 done. That is the payoff of one feature per commit: the culprit is not "somewhere in Tuesday's work" but one named feature, with its own diff, its own verify command and its own event log from the session that built it.

A few practical notes. The test command must be reliable — bisecting with a flaky test gives random answers, which is one more reason to quarantine flaky tests early. Exit code 125 tells bisect to skip a commit that cannot be tested. And if a commit mixed two features, bisect can only point at both; that is why the harness, not the agent, decides when to commit.

Check your understanding

0 of 3 answered

1.A Kite session ends with gate_failed. What state is the repository in afterwards?

2.Fourteen per-feature commits sit between a good and a bad build. About how many test runs does git bisect run need to find the first bad commit?

3.A feature merged last week is wrong. The branch is shared with other engineers. How should it be undone?