Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Reading a failed run: a diagnosis checklist
A failed 60-turn session leaves a log of 2 MB. Reading it from top to bottom, carefully, takes about 45 minutes, and by turn 40 you have forgotten what happened at turn 8. Most people give up and re-run the task, hoping for better luck. Sometimes it works, and you learn nothing.
There is a faster way. Most failed runs fall into a small number of patterns, and each pattern leaves clear marks in the log. If you read the log in the right order — end first, then the shape, then the details — you can usually name the cause in five minutes. This lesson gives you that order as a checklist, a small script that does the first pass for you, and a table of patterns.
Start at the end
The session_end event tells you which stop condition fired, and each one points to a different first question:
| Status | What it means | First question |
|---|---|---|
gate_failed | The agent claimed "done" three times and the gate refused each time | Was it the same check every time? |
out_of_turns | The turn limit was reached | Was the agent making progress, or repeating itself? |
out_of_tokens | The token budget was spent | Which tool results were large, or which files were read again and again? |
cut_off | A reply hit max_tokens, or the model refused | What was the model trying to write at that point? |
No session_end | The harness crashed | What is the last event, and what exception is in the terminal? |
Starting at the end saves you from reading 40 turns of perfectly good work before reaching the part that matters.
The checklist
- Status and totals — read
session_end: status, turns, tokens. Compare with the budgets. A run that ended at turn 60 of 60 is a different story from one that ended at turn 9. - The shape — print the timeline, one line per model call and tool call. Look for repetition: the same command, the same file, the same gate failure.
- Tool errors — count results with
is_errortrue. Is one error repeating? Did the agent change approach after it, or retry the same thing? - Refusals — find results that start with "Not permitted". Was the agent blocked by a rule it needed to break for the task, or trying something it should not do?
- Gate feedback — for each gate event, which check failed? The same one every time suggests missing knowledge, not carelessness.
- The first wrong belief — scan the model's text for the earliest turn where it states something false: "the fee is calculated in
utils.py", "this test is unrelated". Everything after that turn rests on it. - The owner — ask the contractor question from the first section: would a strong human with the same instructions, tools and repository have succeeded? Write one line in the failure log with the layer and the fix.
A first pass in thirty lines
Steps 1 to 5 are mechanical, so a script can do them:
1import json2import sys3from collections import Counter4from pathlib import Path567def diagnose(path: Path) -> list[str]:8 events = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]9 tools = [e for e in events if e["kind"] == "tool"]10 end = next((e for e in events if e["kind"] == "session_end"), {"status": "crashed"})11 notes = [f"ended: {end['status']} after {end.get('turns', '?')} turns"]12 calls = Counter((e["name"], json.dumps(e["args"], sort_keys=True)) for e in tools)13 notes += [f"same call {n} times: {name} {args[:70]}" for (name, args), n in calls.items() if n >= 3]14 errors = [e for e in tools if e["is_error"]]15 notes.append(f"tool errors: {len(errors)} of {len(tools)} calls")16 refused = [e for e in errors if e["output"].startswith("Not permitted")]17 if refused:18 notes.append(f"refused: {len(refused)}, first was {refused[0]['name']} {refused[0]['args']}")19 for i, gate in enumerate((e for e in events if e["kind"] == "gate"), 1):20 reason = "passed" if gate["passed"] else gate["feedback"].split("\n\n")[1].splitlines()[0]21 notes.append(f"gate {i}: {reason[:90]}")22 return notes232425if __name__ == "__main__":26 print("\n".join(diagnose(Path(sys.argv[1]))))The gate reason is taken from the feedback's second paragraph, because Kite's gate message puts each problem in its own paragraph after a fixed first line. On a failed Ledgerly session for F3, the CSV export, it prints:
ended: gate_failed after 31 turnssame call 7 times: run {"command": "python -m pytest -q tests/unit/test_export_csv.py"}tool errors: 1 of 24 callsrefused: 1, first was run {'command': 'pip install freezegun'}gate 1: 'bash scripts/check.sh' exited with code 1:gate 2: 'bash scripts/check.sh' exited with code 1:gate 3: 'bash scripts/check.sh' exited with code 1:Seven runs of the same test, one refused install, and the same gate failure three times. That is already most of the story.
Patterns and their owners
| Sign in the log | Likely pattern | Usual owner and fix |
|---|---|---|
| Same failing command many times, small edits between | Stuck loop | Harness limits stop it; notes and a human hint fix it |
| Same file read 5 or more times | Lost in a large file or large context | Repo (smaller modules, a map) or harness (routes in the task) |
| Several refusals for the same need | Permission wall | Harness policy is too tight, or the task needs a human step first |
| Same gate check fails 3 times | Missing knowledge: a fixture, a rule, a setting | Repo or instructions; tell the agent what it cannot discover |
| Late turns touch files unrelated to the task | Drift | Scope rule in instructions; the gate's scope checks |
| Tool output clipped just where the error was | Bad tool output | Harness: clip differently, or offer a narrower command |
| Right facts in context, wrong conclusion, across runs | Model limit | Stronger model, or split the task |
The table is a starting point, not a verdict. The same sign can have two causes. Seven identical test runs might be a stuck loop, or they might be a flaky test the agent is — reasonably — trying to reproduce. Step 6, finding the first wrong belief, is what separates them.
Walking the F3 example through
With the first pass printed, the rest of the F3 diagnosis took four minutes. The refused call was pip install freezegun: the agent wanted a library to freeze the clock in its date-range tests, and unattended mode refused it, as it should. The agent then tried to control dates by patching datetime.date.today directly. That fails in Python, because the built-in type cannot be patched that way, and it produced the same TypeError seven times with small variations.
The first wrong belief appeared at turn 9: "I'll patch datetime.date.today so the export sees a fixed date." Everything after that turn rested on it. The contractor question: a strong human would have known that trick fails, but might also have asked for freezegun, and been told no. Ledgerly already had the answer — ledgerly/clock.py with a today() function built for exactly this, used by every other date test — but nothing pointed to it. Owner: the repo and instructions. Fix: one route in AGENTS.md: "Dates in tests: use ledgerly.clock, never patch datetime." The next F3 session used it at turn 4 and passed the gate at turn 14.
Check your understanding
0 of 3 answered
1.A session ends out_of_turns. The timeline shows read_file ledgerly/utils.py 11 times, with the agent searching for callers of parse_date. What is the most likely pattern and fix?
2.Why does the checklist ask you to find "the first wrong belief" before deciding the owner?
3.The gate failed three times on the same missing test fixture, which exists only on the CI server. Which layer owns the fix?