Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Why long agent runs go wrong: small slips add up
Watch an agent's first ten minutes on a task and you will usually be impressed. It greps for the right names, reads the right files and makes a sensible first edit. Watch the same run at minute 70 and it can look like a different system. It re-reads files it already read. It "fixes" a test it broke itself. It starts refactoring a module that has nothing to do with the task.
Nothing changed in the model between minute one and minute 70. What changed is the situation the model is in. This lesson names the three mechanisms behind that decay, puts numbers on them, and points to the harness features that push back against each one.
Mechanism 1: the context fills up
The model sees the whole message list on every turn: system prompt, task, and every tool call and result so far. On Ledgerly, an average turn adds about 3,000 tokens — a file read, a test run's output, the model's short reasoning. Start from a 4,000-token system prompt and task, and the context the model reads grows like this:
| Turn | Context read on this turn | Total input tokens so far | Cost so far, uncached |
|---|---|---|---|
| 1 | 4,000 | 4,000 | 0.02 USD |
| 10 | 31,000 | 175,000 | 0.88 USD |
| 30 | 91,000 | 1,425,000 | 7.12 USD |
| 60 | 181,000 | 5,550,000 | 27.75 USD |
The cost column assumes 5 USD per million input tokens, a typical price for a frontier model. Check your provider's current price; the shape matters more than the numbers.
Two things grow here. First, cost grows faster than turns, because each turn re-reads everything before it. Twice the turns costs roughly four times as much. Prompt caching softens this a lot: re-read tokens are billed at about a tenth of the price, so the 60-turn run drops to about 4 USD. But caching does nothing for the second problem.
Second, attention thins out. Even with a context window of a million tokens, the model has to find the three lines that matter among 181,000. The task statement is 180,000 tokens back. The most recent test failure is right at the end, loud and specific. Models, like people, give more weight to what is recent and vivid.
1def run_cost(turns: int, base: int = 4_000, per_turn: int = 3_000,2 usd_per_million: float = 5.0, cached_share: float = 0.0) -> float:3 """Total input cost when every turn re-reads the whole history."""4 total = sum(base + per_turn * (t - 1) for t in range(1, turns + 1))5 fresh = total * (1 - cached_share)6 cached = total * cached_share * 0.1 # cache reads cost about a tenth7 return (fresh + cached) / 1_000_000 * usd_per_million89for turns in (10, 30, 60):10 print(turns, round(run_cost(turns), 2), round(run_cost(turns, cached_share=0.95), 2))Run it and you get about 0.88, 7.12 and 27.75 USD without caching, and about 0.13, 1.03 and 4.02 USD with 95% of input served from cache. Put your own repository's numbers in — measure the average tokens per turn from a real log — and you have a budget you can defend.
Mechanism 2: the goal drifts
Goal drift is the agent slowly swapping the task it was given for the problem in front of it. On Ledgerly it looks like this. The task: add a 2% late fee. At turn 22, the full suite shows the flaky PDF test failing. The agent now has two goals: the late fee, and "make the suite green". The second is more recent and more concrete. By turn 40 the agent is reading PDF renderer code. By turn 55 it has marked the test as skipped "because it is unrelated to the change and flaky" — which is true, and exactly the wrong decision to make alone.
Drift is not disobedience. The model is doing what its context rewards. The fix is to change the context: keep sessions short so the task is never far away, tell the agent in advance what to do with problems it finds (the next section calls this scope control), and have the harness — not the model — decide what counts as green.
Mechanism 3: early mistakes compound
Each step an agent takes has some chance of being wrong. A wrong step early on — a misread function, a wrong assumption about where fees are calculated — becomes the ground every later step stands on. Suppose each step is independently right with probability p. The chance that a whole run has no uncorrected mistake is p multiplied by itself once per step:
| Per-step reliability | 10 steps | 30 steps | 60 steps |
|---|---|---|---|
| 99% | 90% | 74% | 55% |
| 98% | 82% | 55% | 30% |
| 95% | 60% | 22% | 5% |
Real steps are not independent, and a good agent corrects some mistakes itself, so treat this as a picture rather than a prediction. But the picture is right: long runs with no checkpoints fail far more often than short ones, even when each step looks competent. A 98%-reliable agent that sounds impressive in a demo has only a 30% chance of a clean 60-step run.
The harness answer is to break the chain. A test run after an edit is a checkpoint: it catches a wrong step before twenty more are built on it. A completion gate is a checkpoint at the end. A session boundary with a progress file is a checkpoint between chunks of work. Each one resets the multiplication.
One long session
- 60 turns, task statement 180,000 tokens back by the end
- One mistake at turn 8 poisons turns 9 to 60
- About 28 USD uncached, rising with the square of the length
- Nobody knows where it went wrong
Short sessions with checkpoints
- 4 sessions of 15 turns, each starting fresh with the task in view
- A gate at the end of each session catches mistakes early
- About 7.50 USD uncached for the same 60 turns, plus some re-reading at each start
- Each session has its own log and commit
What the harness does about it
Each mechanism has a matching harness feature, and together they are the plan for the rest of the course:
- Full context: turn and token budgets, tools that return short useful output, and sessions sized to one feature.
- Goal drift: instruction files that say what to do with unrelated problems, and a rule that the agent queues discovered work instead of doing it.
- Compounding errors: fast tests in the inner loop, a completion gate at the end, and git commits as save points you can roll back to.
There is a trade-off. Short sessions mean the agent re-reads some files at the start of each one, which costs tokens. On Ledgerly that overhead is about 15,000 tokens per session. That is cheap compared with the cost of turn 60 of a long session, but it is not free, and for tasks that genuinely need one long line of thought — a tricky debugging session, say — splitting too early can hurt.
Check your understanding
0 of 3 answered
1.An agent run costs 1 USD at 20 turns. Roughly what will a 40-turn run of the same kind cost without prompt caching?
2.At turn 30, an agent adding a CSV export notices that utils.parse_date mishandles two date formats and starts fixing it. Which mechanism is this, and what is the harness-level fix?
3.Why does a completion gate help against compounding errors, even though it only runs at the end?