Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Making runs observable
On a Tuesday night, the Ledgerly team left Kite running a batch of five feature sessions. In the morning, three features were done and two had failed. The bill for the night was 38 USD. And one session, F7, had used 2.4 million tokens — four times as many as the others. Why? Nobody could say. The only record of each session was the agent's final summary, and F7's summary said, cheerfully, that it had "made good progress".
You cannot improve what you cannot see. A harness that records nothing turns every failure into guesswork and every cost spike into a mystery. This lesson adds the missing instrument panel: an event log that records every step, and budgets that are set from what the log shows.
One line per event
An event log is a file with one JSON object per line, written as the run happens. JSON Lines is a good format for this: it can be appended to safely, read by any language, searched with grep, and it survives a crash, because every finished line is already on disk.
Kite records five kinds of event:
| Event | Key fields | Answers the question |
|---|---|---|
session_start | feature, model id | What was this run trying to do, with what? |
model | turn, ms, stop reason, content, input and output tokens | What did the model say, how long did it take, what did it cost? |
tool | name, args, ms, is_error, output (clipped) | What did the agent do, and what came back? |
gate | passed, feedback | Did "done" survive verification, and if not, why? |
session_end | status, turns, tokens | How did it end? |
Here are three real lines from a failed Ledgerly session, shortened:
{"t": 0.0, "kind": "session_start", "feature": "F3", "model": "claude-opus-5", "replay": ""}{"t": 4.1, "kind": "model", "turn": 1, "ms": 4080, "stop": "tool_use", "content": [{"type": "tool_use", "id": "toolu_01", "name": "run", "input": {"command": "git push origin main"}}], "input_tokens": 4212, "output_tokens": 61}{"t": 4.1, "kind": "tool", "name": "run", "args": {"command": "git push origin main"}, "ms": 0, "is_error": true, "output": "Not permitted (deny): 'git push origin main' matches the deny rule ..."}The model event stores the model's full content blocks, exactly as returned. That costs disk space — a 30-turn session is typically 1 to 3 MB — and it buys two things. You can read exactly what the model said at every turn. And, as the lesson on testing shows, you can replay the run later, because the log contains every reply the model gave.
The t field is seconds since the session started, from a monotonic clock. Gaps between events tell you where time went: slow model calls, slow tests, or a human taking four minutes to answer an approval prompt.
Logging by wrapping, not by editing the loop
The simplest way to add logging is to scatter log.emit(...) calls through the loop. It works, and it makes the loop harder to read and harder to test. Kite does something cleaner. The loop only knows three interfaces: a model with complete(), a toolbox with specs() and run(), and a gate with check(). So logging is added by wrapping each one in an object with the same interface:
1class LoggedModel:2 def __init__(self, inner, log: EventLog):3 self.inner, self.log, self.turn = inner, log, 045 def complete(self, system, messages, tools):6 self.turn += 17 began = time.monotonic()8 reply = self.inner.complete(system, messages, tools)9 self.log.emit("model", turn=self.turn, ms=int((time.monotonic() - began) * 1000),10 stop=reply.stop_reason, content=reply.content,11 input_tokens=reply.input_tokens, output_tokens=reply.output_tokens)12 return replyLoggedToolbox and LoggedGate follow the same pattern. The loop receives the wrapped objects and never knows it is being watched. That keeps the loop exactly as it was, and it means the same logging works for the real model, the scripted fake used in tests, and a replay.
Budgets from evidence
The first section showed why every run needs a turn limit and a token budget. The event log tells you what those limits should be. Here is a short script that reads a folder of logs and prints one line per session, plus the spread of turns for successful sessions:
1import json2import statistics3import sys4from pathlib import Path56PRICE_IN, PRICE_OUT = 5.0, 25.0 # USD per million tokens; use your provider's prices789def session(path: Path) -> dict:10 events = [json.loads(line) for line in path.read_text().splitlines() if line.strip()]11 models = [e for e in events if e["kind"] == "model"]12 end = next((e for e in events if e["kind"] == "session_end"), {"status": "crashed"})13 tokens_in = sum(e["input_tokens"] for e in models)14 tokens_out = sum(e["output_tokens"] for e in models)15 return {"name": path.stem, "status": end["status"], "turns": len(models),16 "tokens": tokens_in + tokens_out, "minutes": events[-1]["t"] / 60,17 "cost": tokens_in / 1e6 * PRICE_IN + tokens_out / 1e6 * PRICE_OUT,18 "tool_errors": sum(e["is_error"] for e in events if e["kind"] == "tool")}192021rows = [session(p) for p in sorted(Path(sys.argv[1]).glob("*.jsonl"))]22for r in rows:23 print(f"{r['name']:24} {r['status']:13} turns={r['turns']:<3} tokens={r['tokens']:>9,} "24 f"cost<={r['cost']:5.2f} USD {r['minutes']:4.1f} min tool errors={r['tool_errors']}")25done = sorted(r["turns"] for r in rows if r["status"] == "done")26if len(done) >= 2:27 print(f"done: {len(done)} of {len(rows)}; turns p50={statistics.median(done)}, "28 f"p90={statistics.quantiles(done, n=10)[-1]:.0f}")Run it as python costs.py .kite/runs. The cost column is marked with "less than or equal" on purpose. Kite's input token count includes tokens served from the prompt cache, which are billed at a fraction of the price, so this is an upper bound. For exact numbers, log the cache fields separately or use your provider's usage reports.
Over its first 40 successful sessions, Ledgerly's numbers were: median 21 turns, 90th percentile 34 turns, 90th percentile 1.4 million tokens. The rule of thumb is to set each budget at about twice the 90th percentile of successful runs. That gives Kite's defaults: 60 turns and 3 million tokens. A budget set this way almost never stops a session that would have succeeded, and it caps the cost of a stuck session at a number you chose in advance.
Revisit the budgets when the numbers move. A new model, a bigger repository or a new kind of task changes the distribution, and a budget from last quarter can quietly become either useless or too tight.
What not to log, and where logs go
Tool output can contain secrets. If an agent runs env, or a test prints a connection string, the log has it. Three habits keep that manageable. Clip tool output in the log (Kite keeps 2,000 characters per result). Keep logs next to the run, not in a shared bucket: Kite writes a .gitignore containing * into its log folder, so logs never land in a commit by accident. And scrub before you ship logs anywhere else, such as a tracing service.
If your team already uses a tracing system, you can send the same events there as spans: one span per session, child spans per model call and tool call. The OpenTelemetry project publishes conventions for describing model calls. The JSON Lines file is still worth keeping, because you can read it, diff it, and replay it without any service running.
Check your understanding
0 of 3 answered
1.Why does Kite's model event store the model's full content blocks rather than just the token counts?
2.Ledgerly's successful sessions have a median of 21 turns and a 90th percentile of 34 turns. Which turn limit follows the rule of thumb in this lesson?
3.What is the main advantage of adding logging through wrapper objects like LoggedModel instead of adding log calls inside the loop?