Harness Engineering: Making Coding Agents Dependable

Milestone 1: the loop and the message list


In this section you build Kite, the harness from the first lesson, in six milestones. Each milestone adds one part, ends with passing tests, and leaves you with a program that runs. By the end you will have about 450 lines of Python in nine files, a test suite that runs in under ten seconds without spending a token, and a harness you can point at Ledgerly or at your own repository.

The first milestone builds the centre: the message types, the model interface, and the loop. Everything else attaches to these. The most important design decision in the whole build is made here — the loop talks to the model through a single method — because that is what lets every later test use a fake model, and every later feature wrap the model without changing the loop.

You need Python 3.11 or newer. Create a project folder next to your Ledgerly clone, make a virtual environment, and install three packages: python -m venv .venv, activate it, then pip install anthropic pytest ruff.

The message list after one tool calluser: the taskassistant:tool_use c1user:tool_result c1assistant:text, end_turnnullcontentkept verbatimmatching idOutcome: doneEvery exit from run_session returns an Outcome with a named status.
Storing the model's blocks verbatim keeps each tool_use paired with its result, and is what makes replay possible later.

The layout

Here is where every file will end up. You create the first four now.

Text
kite/  __init__.py      empty  config.py        limits and the model id              milestone 1  model.py         messages, Model, adapters            milestone 1  loop.py          the agent loop                       milestone 1, changed once in 4  tools.py         read, write, run                     milestone 2  permissions.py   allow, ask, deny                     milestone 3  __main__.py      the command line                     milestone 3, grows in 4 to 6  gate.py          the completion gate                  milestone 4  progress.py      progress file and save points        milestone 5  events.py        event log and replay                 milestone 6tests/  conftest.py      a tiny Ledgerly for tests            milestone 4  test_*.py        tests for each part

Create an empty kite/__init__.py. Then the configuration, kite/config.py:

Python
import osfrom dataclasses import dataclass, field@dataclassclass Config:    model: str = os.environ.get("KITE_MODEL", "claude-opus-5")    max_tokens: int = 16_000         # ceiling for one model reply    max_turns: int = 60              # model calls in one session    token_budget: int = 3_000_000    # input + output tokens in one session    command_timeout: int = 180       # seconds for one shell command    output_limit: int = 10_000       # characters of tool output the model sees    max_gate_failures: int = 3       # failed "I'm done" claims before giving up    checks: list[str] = field(default_factory=lambda: ["bash scripts/check.sh"])    progress_file: str = "kite-progress.json"    log_dir: str = ".kite/runs"

Every number from the earlier lessons lives here: 60 turns and 3 million tokens from the budgets lesson, three gate attempts, and the one check command from the initialisation lesson. Some fields are used only by later milestones. The model id comes from an environment variable with a default, so you can switch models without touching code.

Message types and the Model interface

kite/model.py starts with the data that flows through the loop:

Python
"""Message types, the Model protocol, a real adapter and a scripted fake."""import copyfrom dataclasses import dataclassfrom typing import Protocol@dataclassclass ToolCall:    id: str    name: str    args: dict@dataclassclass ToolResult:    call_id: str    output: str    is_error: bool = False    def to_block(self) -> dict:        return {"type": "tool_result", "tool_use_id": self.call_id,                "content": self.output, "is_error": self.is_error}@dataclassclass Reply:    content: list[dict]      # the assistant's blocks, kept verbatim for the history    stop_reason: str         # end_turn, tool_use, max_tokens, refusal, ...    input_tokens: int = 0    output_tokens: int = 0    @property    def text(self) -> str:        return "".join(b["text"] for b in self.content if b["type"] == "text")    @property    def tool_calls(self) -> list[ToolCall]:        return [ToolCall(b["id"], b["name"], b["input"])                for b in self.content if b["type"] == "tool_use"]class Model(Protocol):    def complete(self, system: str, messages: list[dict], tools: list[dict]) -> Reply: ...

A Reply stores the model's content blocks as plain dictionaries, exactly as they came back, and derives text and tool_calls from them. Storing the raw blocks follows the rule from the anatomy lesson: the history must hold what the model actually said. It also makes a reply easy to write to a log and read back later, which milestone 6 depends on.

ToolResult.to_block() produces the tool_result block the API expects, with the matching tool_use_id and the is_error flag. The Model protocol is the one-method interface: anything with a matching complete() method is a model as far as Kite is concerned.

Two models: real and scripted

Add the real adapter and the fake to the same file:

Python
class AnthropicModel:    def __init__(self, model: str, max_tokens: int):        import anthropic                     # only the real adapter needs the SDK        self.client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY        self.model, self.max_tokens = model, max_tokens    def complete(self, system, messages, tools):        resp = self.client.messages.create(            model=self.model, max_tokens=self.max_tokens, system=system,            messages=messages, tools=tools,            cache_control={"type": "ephemeral"})  # re-sent history is billed at cache price        u = resp.usage        seen = u.input_tokens + (u.cache_read_input_tokens or 0) + (u.cache_creation_input_tokens or 0)        blocks = [b.model_dump(exclude_none=True) for b in resp.content]        return Reply(blocks, resp.stop_reason, seen, u.output_tokens)class ScriptedModel:    """Plays back prepared replies. Used by tests, and by replay in milestone 6."""    def __init__(self, replies: list[Reply]):        self.replies = list(replies)        self.seen: list[list[dict]] = []     # the message list shown on each call    def complete(self, system, messages, tools):        self.seen.append(copy.deepcopy(messages))        if not self.replies:            raise RuntimeError("ScriptedModel has no replies left")        return self.replies.pop(0)def say(text: str) -> Reply:    return Reply([{"type": "text", "text": text}], "end_turn")def use(name: str, call_id: str = "c1", **args) -> Reply:    return Reply([{"type": "tool_use", "id": call_id, "name": name, "input": args}], "tool_use")

Three details in AnthropicModel matter. cache_control turns on automatic prompt caching for the request, so the growing history is billed at the cache price after the first read, which is the difference between 28 and about 4 USD in the long-run example from the first section. model_dump(exclude_none=True) turns each SDK block into a plain dictionary without dropping fields, including any thinking blocks the model returns, which must be sent back unchanged. And the token count adds cached and uncached input together, because the budget is about how much the model read, not how it was billed.

ScriptedModel records a deep copy of every message list it is shown. The copy matters: the loop keeps appending to the same list, and without a copy every entry in seen would end up showing the final state.

The loop

kite/loop.py is short. It is the 30-line loop from the anatomy lesson with its gaps closed:

Python
"""The agent loop: ask the model, run its tools, repeat until it stops."""from dataclasses import dataclass@dataclassclass Outcome:    status: str    # done, cut_off, out_of_turns, out_of_tokens, gate_failed    turns: int    tokens: int    summary: strdef run_session(task: str, model, toolbox, cfg, *, system: str) -> Outcome:    messages = [{"role": "user", "content": task}]    tokens = 0    for turn in range(1, cfg.max_turns + 1):        reply = model.complete(system, messages, toolbox.specs())        tokens += reply.input_tokens + reply.output_tokens        messages.append({"role": "assistant", "content": reply.content})        if reply.stop_reason in ("max_tokens", "refusal"):            return Outcome("cut_off", turn, tokens, reply.stop_reason)        calls = reply.tool_calls        if not calls:            return Outcome("done", turn, tokens, reply.text)        results = [toolbox.run(call) for call in calls]        messages.append({"role": "user", "content": [r.to_block() for r in results]})        if tokens >= cfg.token_budget:            return Outcome("out_of_tokens", turn, tokens, reply.text)    return Outcome("out_of_turns", cfg.max_turns, tokens, "")

Every exit returns an Outcome with a status from the stop-conditions table in the first section. The loop never raises for a normal stop, so the caller always gets something it can record. Tool results for one turn go back in a single user message. The loop knows nothing about what the tools do: it asks a toolbox for specs() and calls run(). You build the real toolbox in the next milestone. For now, a test can pass in anything with those two methods.

Test it

Create tests/test_loop.py:

Python
from kite.config import Configfrom kite.loop import run_sessionfrom kite.model import ScriptedModel, ToolResult, say, useclass EchoTools:    def specs(self):        return [{"name": "echo", "description": "Echo the text back.",                 "input_schema": {"type": "object", "properties": {"text": {"type": "string"}}}}]    def run(self, call):        return ToolResult(call.id, call.args["text"])def test_tool_result_goes_back_to_the_model():    model = ScriptedModel([use("echo", text="hello"), say("The tool said hello.")])    outcome = run_session("Use the tool.", model, EchoTools(), Config(), system="test")    assert (outcome.status, outcome.turns) == ("done", 2)    shown = model.seen[1]                       # what the model saw on its second call    assert [m["role"] for m in shown] == ["user", "assistant", "user"]    assert shown[2]["content"] == [{"type": "tool_result", "tool_use_id": "c1",                                    "content": "hello", "is_error": False}]def test_turn_limit_stops_a_model_that_never_finishes():    model = ScriptedModel([use("echo", text="again")] * 10)    outcome = run_session("Loop.", model, EchoTools(), Config(max_turns=3), system="test")    assert (outcome.status, outcome.turns) == ("out_of_turns", 3)

Run python -m pytest -q from the project folder. Both tests pass without an API key. The first proves the message list has the right shape on the model's second call. The second proves a model that never stops is stopped. If you want to see a real model go round the loop once, set ANTHROPIC_API_KEY and run run_session with AnthropicModel(cfg.model, cfg.max_tokens) and the same EchoTools, asking it to use the echo tool and report what came back.

Check your understanding

0 of 3 answered

1.Why does ScriptedModel store a deep copy of messages on each call instead of the list itself?

2.A reply comes back with stop_reason equal to max_tokens and no tool calls. What does Kite's loop do?

3.Why does AnthropicModel count cached and uncached input tokens together?