Harness Engineering: Making Coding Agents Dependable

Reading a failed run: a diagnosis checklist


A failed 60-turn session leaves a log of 2 MB. Reading it from top to bottom, carefully, takes about 45 minutes, and by turn 40 you have forgotten what happened at turn 8. Most people give up and re-run the task, hoping for better luck. Sometimes it works, and you learn nothing.

There is a faster way. Most failed runs fall into a small number of patterns, and each pattern leaves clear marks in the log. If you read the log in the right order — end first, then the shape, then the details — you can usually name the cause in five minutes. This lesson gives you that order as a checklist, a small script that does the first pass for you, and a table of patterns.

Read a failed log in this orderStatus and totals — which stop firedShape — repeated calls and filesErrors, refusals, gate feedbackThe first wrong beliefOwner — model, harness or repo
In the failed F3 run, the first wrong belief was patching datetime at turn 9, and every later turn rested on it.

Start at the end

The session_end event tells you which stop condition fired, and each one points to a different first question:

StatusWhat it meansFirst question
gate_failedThe agent claimed "done" three times and the gate refused each timeWas it the same check every time?
out_of_turnsThe turn limit was reachedWas the agent making progress, or repeating itself?
out_of_tokensThe token budget was spentWhich tool results were large, or which files were read again and again?
cut_offA reply hit max_tokens, or the model refusedWhat was the model trying to write at that point?
No session_endThe harness crashedWhat is the last event, and what exception is in the terminal?

Starting at the end saves you from reading 40 turns of perfectly good work before reaching the part that matters.

The checklist

  1. Status and totals — read session_end: status, turns, tokens. Compare with the budgets. A run that ended at turn 60 of 60 is a different story from one that ended at turn 9.
  2. The shape — print the timeline, one line per model call and tool call. Look for repetition: the same command, the same file, the same gate failure.
  3. Tool errors — count results with is_error true. Is one error repeating? Did the agent change approach after it, or retry the same thing?
  4. Refusals — find results that start with "Not permitted". Was the agent blocked by a rule it needed to break for the task, or trying something it should not do?
  5. Gate feedback — for each gate event, which check failed? The same one every time suggests missing knowledge, not carelessness.
  6. The first wrong belief — scan the model's text for the earliest turn where it states something false: "the fee is calculated in utils.py", "this test is unrelated". Everything after that turn rests on it.
  7. The owner — ask the contractor question from the first section: would a strong human with the same instructions, tools and repository have succeeded? Write one line in the failure log with the layer and the fix.

A first pass in thirty lines

Steps 1 to 5 are mechanical, so a script can do them:

Python
import jsonimport sysfrom collections import Counterfrom pathlib import Pathdef diagnose(path: Path) -> list[str]:    events = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]    tools = [e for e in events if e["kind"] == "tool"]    end = next((e for e in events if e["kind"] == "session_end"), {"status": "crashed"})    notes = [f"ended: {end['status']} after {end.get('turns', '?')} turns"]    calls = Counter((e["name"], json.dumps(e["args"], sort_keys=True)) for e in tools)    notes += [f"same call {n} times: {name} {args[:70]}" for (name, args), n in calls.items() if n >= 3]    errors = [e for e in tools if e["is_error"]]    notes.append(f"tool errors: {len(errors)} of {len(tools)} calls")    refused = [e for e in errors if e["output"].startswith("Not permitted")]    if refused:        notes.append(f"refused: {len(refused)}, first was {refused[0]['name']} {refused[0]['args']}")    for i, gate in enumerate((e for e in events if e["kind"] == "gate"), 1):        reason = "passed" if gate["passed"] else gate["feedback"].split("\n\n")[1].splitlines()[0]        notes.append(f"gate {i}: {reason[:90]}")    return notesif __name__ == "__main__":    print("\n".join(diagnose(Path(sys.argv[1]))))

The gate reason is taken from the feedback's second paragraph, because Kite's gate message puts each problem in its own paragraph after a fixed first line. On a failed Ledgerly session for F3, the CSV export, it prints:

Text
ended: gate_failed after 31 turnssame call 7 times: run {"command": "python -m pytest -q tests/unit/test_export_csv.py"}tool errors: 1 of 24 callsrefused: 1, first was run {'command': 'pip install freezegun'}gate 1: 'bash scripts/check.sh' exited with code 1:gate 2: 'bash scripts/check.sh' exited with code 1:gate 3: 'bash scripts/check.sh' exited with code 1:

Seven runs of the same test, one refused install, and the same gate failure three times. That is already most of the story.

Patterns and their owners

Sign in the logLikely patternUsual owner and fix
Same failing command many times, small edits betweenStuck loopHarness limits stop it; notes and a human hint fix it
Same file read 5 or more timesLost in a large file or large contextRepo (smaller modules, a map) or harness (routes in the task)
Several refusals for the same needPermission wallHarness policy is too tight, or the task needs a human step first
Same gate check fails 3 timesMissing knowledge: a fixture, a rule, a settingRepo or instructions; tell the agent what it cannot discover
Late turns touch files unrelated to the taskDriftScope rule in instructions; the gate's scope checks
Tool output clipped just where the error wasBad tool outputHarness: clip differently, or offer a narrower command
Right facts in context, wrong conclusion, across runsModel limitStronger model, or split the task

The table is a starting point, not a verdict. The same sign can have two causes. Seven identical test runs might be a stuck loop, or they might be a flaky test the agent is — reasonably — trying to reproduce. Step 6, finding the first wrong belief, is what separates them.

Walking the F3 example through

With the first pass printed, the rest of the F3 diagnosis took four minutes. The refused call was pip install freezegun: the agent wanted a library to freeze the clock in its date-range tests, and unattended mode refused it, as it should. The agent then tried to control dates by patching datetime.date.today directly. That fails in Python, because the built-in type cannot be patched that way, and it produced the same TypeError seven times with small variations.

The first wrong belief appeared at turn 9: "I'll patch datetime.date.today so the export sees a fixed date." Everything after that turn rested on it. The contractor question: a strong human would have known that trick fails, but might also have asked for freezegun, and been told no. Ledgerly already had the answer — ledgerly/clock.py with a today() function built for exactly this, used by every other date test — but nothing pointed to it. Owner: the repo and instructions. Fix: one route in AGENTS.md: "Dates in tests: use ledgerly.clock, never patch datetime." The next F3 session used it at turn 4 and passed the gate at turn 14.

Check your understanding

0 of 3 answered

1.A session ends out_of_turns. The timeline shows read_file ledgerly/utils.py 11 times, with the agent searching for callers of parse_date. What is the most likely pattern and fix?

2.Why does the checklist ask you to find "the first wrong belief" before deciding the owner?

3.The gate failed three times on the same missing test fixture, which exists only on the CI server. Which layer owns the fix?