Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Anatomy of an agent loop
Strip away the product names and every coding agent — Claude Code, Cursor's agent, Codex, the one you will build — runs the same loop. The model reads a list of messages and either answers or asks to use a tool. The harness runs the tool and appends the result to the list. Repeat.
Everything else in this course is attached to this loop: permissions sit between "asks to use a tool" and "runs the tool", the completion gate sits at "answers", and the event log watches every step. So it is worth knowing the loop precisely — its data, its rules and its exits — before adding anything to it.
The message list is the agent's memory
A model call is stateless. The model remembers nothing between calls. Everything it knows about the task, the files it read and the tests it ran is in the message list you send, every single time. That list alternates between two roles:
- user messages: the task, and later the results of tool calls.
- assistant messages: the model's output, a list of content blocks —
textblocks andtool_useblocks.
After one tool call on Ledgerly, the list looks like this:
1[2 {"role": "user", "content": "How many tests does tests/test_tax.py define? Do not change files."},3 {"role": "assistant", "content": [4 {"type": "text", "text": "I'll count the test functions."},5 {"type": "tool_use", "id": "toolu_01", "name": "run",6 "input": {"command": "grep -c '^def test_' tests/test_tax.py"}}7 ]},8 {"role": "user", "content": [9 {"type": "tool_result", "tool_use_id": "toolu_01", "content": "exit code 0\n17\n"}10 ]}11]Three rules keep this list valid, and breaking any of them gets you an API error or a confused model:
- Every
tool_useblock must be answered by atool_resultblock with the same id, in the very next user message. - When the model asks for several tools at once, all their results go back together in one user message.
- A tool that fails still gets a result: send the error text with
"is_error": true. Never drop it.
How the list grows, and where the tokens go
Here is how the list grew in the F1 late-fee session from the first lesson, read from Kite's event log. "Input" is everything the model read on that turn.
| Turn | What the list gained just before this turn | Input this turn |
|---|---|---|
| 1 | Nothing yet: system prompt, instruction file, three tool definitions, the task | 4,120 |
| 2 | grep -rn due_date: 118 matching lines, about 2,600 tokens | 6,890 |
| 6 | ledgerly/utils.py, lines 1 to 400, about 5,900 tokens | 21,500 |
| 12 | The first failing unit test run, clipped, about 2,500 tokens | 44,800 |
| 19 | The fourth unit test run; the model then says it is finished | 71,340 |
That is about 3,700 tokens a turn, a little above the 3,000 average from the previous lesson, because this session ran the unit tests four times. Here is where the 71,340 tokens at turn 19 came from:
| Source | Tokens | Share |
|---|---|---|
| Fixed prefix: system prompt, instruction file, tool definitions, task | 4,120 | 6% |
| The model's own messages: short reasoning and tool calls, including file contents it wrote | 5,320 | 7% |
| Tool results: file reads | 38,200 | 54% |
| Tool results: test runs | 17,900 | 25% |
Tool results: grep, ls and git | 5,800 | 8% |
Tool results are 87% of the context; the instructions are 6%. And 9,400 of the file-read tokens are pages of utils.py the agent had already read once.
When a token arrives matters too, because everything in the list is read again on every later turn. An unclipped 9,000-token test run at turn 6 of a 24-turn session is read 18 more times: 162,000 input tokens, fifteen times what the 450-token instruction file costs over the whole session. So cut tool results first (clip output, page large files, stop repeated reads) and trim the system prompt last. Caching makes re-reads cheaper, not smaller: the model still has to find the task among 71,000 tokens.
A loop in 30 lines
Here is a complete, working agent with one tool — a shell — using the Anthropic Python SDK. It is deliberately naive; the rest of the course fixes its weaknesses.
1import subprocess23import anthropic45client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY6MODEL = "claude-opus-5" # keep the model id in one place7TOOLS = [{"name": "run", "description": "Run a shell command in the repo; returns exit code and output.",8 "input_schema": {"type": "object", "properties": {"command": {"type": "string"}},9 "required": ["command"]}}]101112def run(command: str) -> str:13 proc = subprocess.run(command, shell=True, cwd="../ledgerly", capture_output=True,14 text=True, timeout=120)15 return f"exit code {proc.returncode}\n{(proc.stdout + proc.stderr)[-8000:]}"161718messages = [{"role": "user", "content": "How many tests does tests/test_tax.py define? Do not change files."}]19for turn in range(1, 21): # stop condition: at most 20 turns20 resp = client.messages.create(model=MODEL, max_tokens=16000, tools=TOOLS,21 system="You are a careful coding agent.", messages=messages)22 messages.append({"role": "assistant", "content": resp.content})23 if resp.stop_reason != "tool_use": # end_turn, max_tokens, refusal24 break25 results = [{"type": "tool_result", "tool_use_id": block.id, "content": run(**block.input)}26 for block in resp.content if block.type == "tool_use"]27 messages.append({"role": "user", "content": results})2829print(f"stopped: {resp.stop_reason} after {turn} turns")30print("".join(block.text for block in resp.content if block.type == "text"))Walk through one turn. The harness sends the whole list plus the tool definitions. The response's stop_reason says why the model stopped writing: tool_use means "run these tools and come back"; anything else means this turn is the last. The harness appends the assistant's content as-is, runs each requested tool, and appends all results in one user message. Then it goes round again.
Parallel tool calls
Models often put several tool_use blocks in one reply, for example to read three files at once. At turn 4 of F1 the context was about 15,300 tokens, so doing three 2,000-token reads in separate turns would cost two extra calls and about 37,000 more input tokens. Parallel calls also raise three questions the harness must answer on purpose:
- Order. A reply can hold an edit and a test run. Run them at the same time and the tests may run before the edit lands. Kite runs calls one after another, in order; run them concurrently only if all are read-only.
- Partial refusal. If the permission layer refuses one call of three, run the other two and return all three results together, the refused one as an error.
- What a turn means. A turn cap counts replies, not tool calls, so log calls per turn too. A "20-turn" session can hide 60 commands.
If your tools cannot be batched safely, set disable_parallel_tool_use in tool_choice and the model makes at most one call per reply, at the price of more turns.
Now look at what this loop does not do. It keeps only the last 8,000 characters of output, which is better than nothing but can cut a traceback in half. A command that runs longer than 120 seconds raises TimeoutExpired and crashes the whole agent. If the model sends a misspelled argument, run(**block.input) raises TypeError and crashes it too. There is no permission check, so git push --force runs like anything else. And "done" is whatever the model says. Every one of these is a harness gap, and each has a lesson.
Stop conditions belong to the harness
The loop above ends in one of three ways. The model decides to stop, the turn limit is reached, or something crashes. A real harness needs a deliberate list:
| Stop condition | Who decides | What the harness should do |
|---|---|---|
end_turn — the model says it is finished | Model | Treat it as a claim; run the completion gate |
max_tokens — the reply hit its length limit | API | Stop and record "cut off"; the reply may be half a tool call |
refusal — the model declined | Model | Stop and record it; do not loop and retry blindly |
| Turn limit reached | Harness | Stop, record "out of turns", keep the work for review |
| Token budget spent | Harness | Stop, record "out of tokens" |
| Gate failed too many times | Harness | Stop, record "gate failed", keep the work out of the main branch |
Notice that only the first row is the model's judgement about the task, and even that one gets checked. The rest are limits the harness enforces, because a model that is stuck will not reliably notice it is stuck.
Budgets that fire at the wrong time
Because every turn re-reads the growing list, total input grows with the square of the turns. At F1's rate of about 3,700 tokens a turn, Kite's 3-million-token budget runs out at turn 40, long before the 60-turn cap; turns 31 to 40 alone read more than turns 1 to 25. The turn cap catches the other kind of stuck session: many small turns. Friday's F2 session in the memory section hit 60 turns having read only 2.4 million tokens, about 1,200 new tokens a turn, while it re-ran one test again and again. Keep both limits; each catches a failure the other misses. And a budget checked after a call can overshoot by that whole call, about 150,000 tokens at turn 40, so if it is a hard ceiling, check the next request's size before sending it.
When "no tool call" means "done"
Both loops in this course treat a reply with no tool calls as "finished". In unattended runs that misfires in three known ways:
- The model announces instead of acting: "Next I'll update the PDF template to show the fee." No call follows, so the loop stops, and the gate may pass because the unfinished part has no test yet.
- The model asks a question ("Should the fee apply to partly paid invoices?") that nobody is there to answer.
- The reply was paused. With server-side tools such as web search, the API can return
pause_turn, which means "send this back so I can continue". A loop that only looks for tool calls stops halfway.
The fixes are small. Check stop_reason in a fixed order: max_tokens and refusal, then pause_turn, then tool calls. When a final reply asks a question or promises a next step, send one nudge: "No one can answer questions in this run. If the task is finished, say so and summarise; if not, continue and record your assumptions." Allow at most two nudges and log each one.
Check your understanding
0 of 3 answered
1.The model returns two tool_use blocks in one reply: read ledgerly/fees.py and run the unit tests. How should the harness send the results back?
2.In the 30-line loop, the model asks to run a command that hangs for five minutes. What happens?
3.Why should a stop_reason of max_tokens end the session rather than be treated like a normal finished answer?