Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Milestone 1: the loop and the message list
In this section you build Kite, the harness from the first lesson, in six milestones. Each milestone adds one part, ends with passing tests, and leaves you with a program that runs. By the end you will have about 450 lines of Python in nine files, a test suite that runs in under ten seconds without spending a token, and a harness you can point at Ledgerly or at your own repository.
The first milestone builds the centre: the message types, the model interface, and the loop. Everything else attaches to these. The most important design decision in the whole build is made here — the loop talks to the model through a single method — because that is what lets every later test use a fake model, and every later feature wrap the model without changing the loop.
You need Python 3.11 or newer. Create a project folder next to your Ledgerly clone, make a virtual environment, and install three packages: python -m venv .venv, activate it, then pip install anthropic pytest ruff.
The layout
Here is where every file will end up. You create the first four now.
kite/ __init__.py empty config.py limits and the model id milestone 1 model.py messages, Model, adapters milestone 1 loop.py the agent loop milestone 1, changed once in 4 tools.py read, write, run milestone 2 permissions.py allow, ask, deny milestone 3 __main__.py the command line milestone 3, grows in 4 to 6 gate.py the completion gate milestone 4 progress.py progress file and save points milestone 5 events.py event log and replay milestone 6tests/ conftest.py a tiny Ledgerly for tests milestone 4 test_*.py tests for each partCreate an empty kite/__init__.py. Then the configuration, kite/config.py:
1import os2from dataclasses import dataclass, field345@dataclass6class Config:7 model: str = os.environ.get("KITE_MODEL", "claude-opus-5")8 max_tokens: int = 16_000 # ceiling for one model reply9 max_turns: int = 60 # model calls in one session10 token_budget: int = 3_000_000 # input + output tokens in one session11 command_timeout: int = 180 # seconds for one shell command12 output_limit: int = 10_000 # characters of tool output the model sees13 max_gate_failures: int = 3 # failed "I'm done" claims before giving up14 checks: list[str] = field(default_factory=lambda: ["bash scripts/check.sh"])15 progress_file: str = "kite-progress.json"16 log_dir: str = ".kite/runs"Every number from the earlier lessons lives here: 60 turns and 3 million tokens from the budgets lesson, three gate attempts, and the one check command from the initialisation lesson. Some fields are used only by later milestones. The model id comes from an environment variable with a default, so you can switch models without touching code.
Message types and the Model interface
kite/model.py starts with the data that flows through the loop:
1"""Message types, the Model protocol, a real adapter and a scripted fake."""2import copy3from dataclasses import dataclass4from typing import Protocol567@dataclass8class ToolCall:9 id: str10 name: str11 args: dict121314@dataclass15class ToolResult:16 call_id: str17 output: str18 is_error: bool = False1920 def to_block(self) -> dict:21 return {"type": "tool_result", "tool_use_id": self.call_id,22 "content": self.output, "is_error": self.is_error}232425@dataclass26class Reply:27 content: list[dict] # the assistant's blocks, kept verbatim for the history28 stop_reason: str # end_turn, tool_use, max_tokens, refusal, ...29 input_tokens: int = 030 output_tokens: int = 03132 @property33 def text(self) -> str:34 return "".join(b["text"] for b in self.content if b["type"] == "text")3536 @property37 def tool_calls(self) -> list[ToolCall]:38 return [ToolCall(b["id"], b["name"], b["input"])39 for b in self.content if b["type"] == "tool_use"]404142class Model(Protocol):43 def complete(self, system: str, messages: list[dict], tools: list[dict]) -> Reply: ...A Reply stores the model's content blocks as plain dictionaries, exactly as they came back, and derives text and tool_calls from them. Storing the raw blocks follows the rule from the anatomy lesson: the history must hold what the model actually said. It also makes a reply easy to write to a log and read back later, which milestone 6 depends on.
ToolResult.to_block() produces the tool_result block the API expects, with the matching tool_use_id and the is_error flag. The Model protocol is the one-method interface: anything with a matching complete() method is a model as far as Kite is concerned.
Two models: real and scripted
Add the real adapter and the fake to the same file:
1class AnthropicModel:2 def __init__(self, model: str, max_tokens: int):3 import anthropic # only the real adapter needs the SDK4 self.client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY5 self.model, self.max_tokens = model, max_tokens67 def complete(self, system, messages, tools):8 resp = self.client.messages.create(9 model=self.model, max_tokens=self.max_tokens, system=system,10 messages=messages, tools=tools,11 cache_control={"type": "ephemeral"}) # re-sent history is billed at cache price12 u = resp.usage13 seen = u.input_tokens + (u.cache_read_input_tokens or 0) + (u.cache_creation_input_tokens or 0)14 blocks = [b.model_dump(exclude_none=True) for b in resp.content]15 return Reply(blocks, resp.stop_reason, seen, u.output_tokens)161718class ScriptedModel:19 """Plays back prepared replies. Used by tests, and by replay in milestone 6."""2021 def __init__(self, replies: list[Reply]):22 self.replies = list(replies)23 self.seen: list[list[dict]] = [] # the message list shown on each call2425 def complete(self, system, messages, tools):26 self.seen.append(copy.deepcopy(messages))27 if not self.replies:28 raise RuntimeError("ScriptedModel has no replies left")29 return self.replies.pop(0)303132def say(text: str) -> Reply:33 return Reply([{"type": "text", "text": text}], "end_turn")343536def use(name: str, call_id: str = "c1", **args) -> Reply:37 return Reply([{"type": "tool_use", "id": call_id, "name": name, "input": args}], "tool_use")Three details in AnthropicModel matter. cache_control turns on automatic prompt caching for the request, so the growing history is billed at the cache price after the first read, which is the difference between 28 and about 4 USD in the long-run example from the first section. model_dump(exclude_none=True) turns each SDK block into a plain dictionary without dropping fields, including any thinking blocks the model returns, which must be sent back unchanged. And the token count adds cached and uncached input together, because the budget is about how much the model read, not how it was billed.
ScriptedModel records a deep copy of every message list it is shown. The copy matters: the loop keeps appending to the same list, and without a copy every entry in seen would end up showing the final state.
The loop
kite/loop.py is short. It is the 30-line loop from the anatomy lesson with its gaps closed:
1"""The agent loop: ask the model, run its tools, repeat until it stops."""2from dataclasses import dataclass345@dataclass6class Outcome:7 status: str # done, cut_off, out_of_turns, out_of_tokens, gate_failed8 turns: int9 tokens: int10 summary: str111213def run_session(task: str, model, toolbox, cfg, *, system: str) -> Outcome:14 messages = [{"role": "user", "content": task}]15 tokens = 016 for turn in range(1, cfg.max_turns + 1):17 reply = model.complete(system, messages, toolbox.specs())18 tokens += reply.input_tokens + reply.output_tokens19 messages.append({"role": "assistant", "content": reply.content})20 if reply.stop_reason in ("max_tokens", "refusal"):21 return Outcome("cut_off", turn, tokens, reply.stop_reason)22 calls = reply.tool_calls23 if not calls:24 return Outcome("done", turn, tokens, reply.text)25 results = [toolbox.run(call) for call in calls]26 messages.append({"role": "user", "content": [r.to_block() for r in results]})27 if tokens >= cfg.token_budget:28 return Outcome("out_of_tokens", turn, tokens, reply.text)29 return Outcome("out_of_turns", cfg.max_turns, tokens, "")Every exit returns an Outcome with a status from the stop-conditions table in the first section. The loop never raises for a normal stop, so the caller always gets something it can record. Tool results for one turn go back in a single user message. The loop knows nothing about what the tools do: it asks a toolbox for specs() and calls run(). You build the real toolbox in the next milestone. For now, a test can pass in anything with those two methods.
Test it
Create tests/test_loop.py:
1from kite.config import Config2from kite.loop import run_session3from kite.model import ScriptedModel, ToolResult, say, use456class EchoTools:7 def specs(self):8 return [{"name": "echo", "description": "Echo the text back.",9 "input_schema": {"type": "object", "properties": {"text": {"type": "string"}}}}]1011 def run(self, call):12 return ToolResult(call.id, call.args["text"])131415def test_tool_result_goes_back_to_the_model():16 model = ScriptedModel([use("echo", text="hello"), say("The tool said hello.")])17 outcome = run_session("Use the tool.", model, EchoTools(), Config(), system="test")18 assert (outcome.status, outcome.turns) == ("done", 2)19 shown = model.seen[1] # what the model saw on its second call20 assert [m["role"] for m in shown] == ["user", "assistant", "user"]21 assert shown[2]["content"] == [{"type": "tool_result", "tool_use_id": "c1",22 "content": "hello", "is_error": False}]232425def test_turn_limit_stops_a_model_that_never_finishes():26 model = ScriptedModel([use("echo", text="again")] * 10)27 outcome = run_session("Loop.", model, EchoTools(), Config(max_turns=3), system="test")28 assert (outcome.status, outcome.turns) == ("out_of_turns", 3)Run python -m pytest -q from the project folder. Both tests pass without an API key. The first proves the message list has the right shape on the model's second call. The second proves a model that never stops is stopped. If you want to see a real model go round the loop once, set ANTHROPIC_API_KEY and run run_session with AnthropicModel(cfg.model, cfg.max_tokens) and the same EchoTools, asking it to use the echo tool and report what came back.
Check your understanding
0 of 3 answered
1.Why does ScriptedModel store a deep copy of messages on each call instead of the list itself?
2.A reply comes back with stop_reason equal to max_tokens and no tool calls. What does Kite's loop do?
3.Why does AnthropicModel count cached and uncached input tokens together?