Harness Engineering: Making Coding Agents Dependable

Anatomy of an agent loop


Strip away the product names and every coding agent — Claude Code, Cursor's agent, Codex, the one you will build — runs the same loop. The model reads a list of messages and either answers or asks to use a tool. The harness runs the tool and appends the result to the list. Repeat.

Everything else in this course is attached to this loop: permissions sit between "asks to use a tool" and "runs the tool", the completion gate sits at "answers", and the event log watches every step. So it is worth knowing the loop precisely — its data, its rules and its exits — before adding anything to it.

One turn of the loopmessage list sentmodel repliestool_use blocksharnessruns toolstool_resultappendedstop_reason decidespermissions go heresame ids,one messageExits: end_turn goes to the gate; turn and token limits belong to the harness.
The model only chooses the next action; the harness decides what runs, what the model sees, and when the loop ends.

The message list is the agent's memory

A model call is stateless. The model remembers nothing between calls. Everything it knows about the task, the files it read and the tests it ran is in the message list you send, every single time. That list alternates between two roles:

  • user messages: the task, and later the results of tool calls.
  • assistant messages: the model's output, a list of content blocks — text blocks and tool_use blocks.

After one tool call on Ledgerly, the list looks like this:

JSON
[  {"role": "user", "content": "How many tests does tests/test_tax.py define? Do not change files."},  {"role": "assistant", "content": [    {"type": "text", "text": "I'll count the test functions."},    {"type": "tool_use", "id": "toolu_01", "name": "run",     "input": {"command": "grep -c '^def test_' tests/test_tax.py"}}  ]},  {"role": "user", "content": [    {"type": "tool_result", "tool_use_id": "toolu_01", "content": "exit code 0\n17\n"}  ]}]

Three rules keep this list valid, and breaking any of them gets you an API error or a confused model:

  1. Every tool_use block must be answered by a tool_result block with the same id, in the very next user message.
  2. When the model asks for several tools at once, all their results go back together in one user message.
  3. A tool that fails still gets a result: send the error text with "is_error": true. Never drop it.

How the list grows, and where the tokens go

Here is how the list grew in the F1 late-fee session from the first lesson, read from Kite's event log. "Input" is everything the model read on that turn.

TurnWhat the list gained just before this turnInput this turn
1Nothing yet: system prompt, instruction file, three tool definitions, the task4,120
2grep -rn due_date: 118 matching lines, about 2,600 tokens6,890
6ledgerly/utils.py, lines 1 to 400, about 5,900 tokens21,500
12The first failing unit test run, clipped, about 2,500 tokens44,800
19The fourth unit test run; the model then says it is finished71,340

That is about 3,700 tokens a turn, a little above the 3,000 average from the previous lesson, because this session ran the unit tests four times. Here is where the 71,340 tokens at turn 19 came from:

SourceTokensShare
Fixed prefix: system prompt, instruction file, tool definitions, task4,1206%
The model's own messages: short reasoning and tool calls, including file contents it wrote5,3207%
Tool results: file reads38,20054%
Tool results: test runs17,90025%
Tool results: grep, ls and git5,8008%

Tool results are 87% of the context; the instructions are 6%. And 9,400 of the file-read tokens are pages of utils.py the agent had already read once.

When a token arrives matters too, because everything in the list is read again on every later turn. An unclipped 9,000-token test run at turn 6 of a 24-turn session is read 18 more times: 162,000 input tokens, fifteen times what the 450-token instruction file costs over the whole session. So cut tool results first (clip output, page large files, stop repeated reads) and trim the system prompt last. Caching makes re-reads cheaper, not smaller: the model still has to find the task among 71,000 tokens.

A loop in 30 lines

Here is a complete, working agent with one tool — a shell — using the Anthropic Python SDK. It is deliberately naive; the rest of the course fixes its weaknesses.

Python
import subprocessimport anthropicclient = anthropic.Anthropic()          # reads ANTHROPIC_API_KEYMODEL = "claude-opus-5"                 # keep the model id in one placeTOOLS = [{"name": "run", "description": "Run a shell command in the repo; returns exit code and output.",          "input_schema": {"type": "object", "properties": {"command": {"type": "string"}},                           "required": ["command"]}}]def run(command: str) -> str:    proc = subprocess.run(command, shell=True, cwd="../ledgerly", capture_output=True,                          text=True, timeout=120)    return f"exit code {proc.returncode}\n{(proc.stdout + proc.stderr)[-8000:]}"messages = [{"role": "user", "content": "How many tests does tests/test_tax.py define? Do not change files."}]for turn in range(1, 21):                                 # stop condition: at most 20 turns    resp = client.messages.create(model=MODEL, max_tokens=16000, tools=TOOLS,                                  system="You are a careful coding agent.", messages=messages)    messages.append({"role": "assistant", "content": resp.content})    if resp.stop_reason != "tool_use":                    # end_turn, max_tokens, refusal        break    results = [{"type": "tool_result", "tool_use_id": block.id, "content": run(**block.input)}               for block in resp.content if block.type == "tool_use"]    messages.append({"role": "user", "content": results})print(f"stopped: {resp.stop_reason} after {turn} turns")print("".join(block.text for block in resp.content if block.type == "text"))

Walk through one turn. The harness sends the whole list plus the tool definitions. The response's stop_reason says why the model stopped writing: tool_use means "run these tools and come back"; anything else means this turn is the last. The harness appends the assistant's content as-is, runs each requested tool, and appends all results in one user message. Then it goes round again.

Parallel tool calls

Models often put several tool_use blocks in one reply, for example to read three files at once. At turn 4 of F1 the context was about 15,300 tokens, so doing three 2,000-token reads in separate turns would cost two extra calls and about 37,000 more input tokens. Parallel calls also raise three questions the harness must answer on purpose:

  • Order. A reply can hold an edit and a test run. Run them at the same time and the tests may run before the edit lands. Kite runs calls one after another, in order; run them concurrently only if all are read-only.
  • Partial refusal. If the permission layer refuses one call of three, run the other two and return all three results together, the refused one as an error.
  • What a turn means. A turn cap counts replies, not tool calls, so log calls per turn too. A "20-turn" session can hide 60 commands.

If your tools cannot be batched safely, set disable_parallel_tool_use in tool_choice and the model makes at most one call per reply, at the price of more turns.

Now look at what this loop does not do. It keeps only the last 8,000 characters of output, which is better than nothing but can cut a traceback in half. A command that runs longer than 120 seconds raises TimeoutExpired and crashes the whole agent. If the model sends a misspelled argument, run(**block.input) raises TypeError and crashes it too. There is no permission check, so git push --force runs like anything else. And "done" is whatever the model says. Every one of these is a harness gap, and each has a lesson.

Stop conditions belong to the harness

The loop above ends in one of three ways. The model decides to stop, the turn limit is reached, or something crashes. A real harness needs a deliberate list:

Stop conditionWho decidesWhat the harness should do
end_turn — the model says it is finishedModelTreat it as a claim; run the completion gate
max_tokens — the reply hit its length limitAPIStop and record "cut off"; the reply may be half a tool call
refusal — the model declinedModelStop and record it; do not loop and retry blindly
Turn limit reachedHarnessStop, record "out of turns", keep the work for review
Token budget spentHarnessStop, record "out of tokens"
Gate failed too many timesHarnessStop, record "gate failed", keep the work out of the main branch

Notice that only the first row is the model's judgement about the task, and even that one gets checked. The rest are limits the harness enforces, because a model that is stuck will not reliably notice it is stuck.

Budgets that fire at the wrong time

Because every turn re-reads the growing list, total input grows with the square of the turns. At F1's rate of about 3,700 tokens a turn, Kite's 3-million-token budget runs out at turn 40, long before the 60-turn cap; turns 31 to 40 alone read more than turns 1 to 25. The turn cap catches the other kind of stuck session: many small turns. Friday's F2 session in the memory section hit 60 turns having read only 2.4 million tokens, about 1,200 new tokens a turn, while it re-ran one test again and again. Keep both limits; each catches a failure the other misses. And a budget checked after a call can overshoot by that whole call, about 150,000 tokens at turn 40, so if it is a hard ceiling, check the next request's size before sending it.

When "no tool call" means "done"

Both loops in this course treat a reply with no tool calls as "finished". In unattended runs that misfires in three known ways:

  • The model announces instead of acting: "Next I'll update the PDF template to show the fee." No call follows, so the loop stops, and the gate may pass because the unfinished part has no test yet.
  • The model asks a question ("Should the fee apply to partly paid invoices?") that nobody is there to answer.
  • The reply was paused. With server-side tools such as web search, the API can return pause_turn, which means "send this back so I can continue". A loop that only looks for tool calls stops halfway.

The fixes are small. Check stop_reason in a fixed order: max_tokens and refusal, then pause_turn, then tool calls. When a final reply asks a question or promises a next step, send one nudge: "No one can answer questions in this run. If the task is finished, say so and summarise; if not, continue and record your assumptions." Allow at most two nudges and log each one.

Check your understanding

0 of 3 answered

1.The model returns two tool_use blocks in one reply: read ledgerly/fees.py and run the unit tests. How should the harness send the results back?

2.In the 30-line loop, the model asks to run a command that hangs for five minutes. What happens?

3.Why should a stop_reason of max_tokens end the session rather than be treated like a normal finished answer?