Harness Engineering: Making Coding Agents Dependable

Milestone 2: tools the model can use


A tool is a promise to the model: "call me with these arguments and I will give you something useful". Most agent failures that look like model confusion are really broken promises. A file read that returns 3,000 lines at once. A test run whose output is cut exactly where the error was. A misspelled argument that crashes the whole agent instead of coming back as a message.

This milestone builds kite/tools.py: three tools — read a file, write a file, run a command — designed around one question: when the model reads this result, can it decide what to do next?

Three tools is deliberately few. With run, the agent can already grep, ls, run tests and use git. Each extra tool is one more description in every request and one more choice for the model. Add tools when a specific failure shows you need one, as the exercise at the end does.

What clip keeps from a huge test runcommandfirstfailurecharacterscutrooterror214errorsin 12.3s01234head, a quartermarkersays how muchtail,three quartersA non-zero exit code is a normal result; is_error means the tool itself failed.
Keeping mostly the end of the output saved the line that mattered from a 38,000-line traceback.

Tool definitions are prompts

The model never sees your Python functions. It sees a name, a description and a JSON schema for each tool, and it decides from those alone. Write them the way you would brief a colleague:

Python
"""Three tools - read, write, run - with results written for a model to read."""import subprocessfrom pathlib import Pathfrom .model import ToolCall, ToolResultSPECS = [    {"name": "read_file",     "description": "Read a text file in the repo. Returns numbered lines, 400 at a time; "                    "pass start to read further down a long file.",     "input_schema": {"type": "object", "properties": {         "path": {"type": "string", "description": "Path relative to the repo root"},         "start": {"type": "integer", "description": "First line to show, 1-based"}},         "required": ["path"]}},    {"name": "write_file",     "description": "Create a file or replace it completely. Send the whole new content.",     "input_schema": {"type": "object", "properties": {         "path": {"type": "string"}, "content": {"type": "string"}},         "required": ["path", "content"]}},    {"name": "run",     "description": "Run a shell command at the repo root, for example "                    "'python -m pytest -q tests/test_tax.py' or 'grep -rn round_money ledgerly'. "                    "Returns the exit code and the end of the output.",     "input_schema": {"type": "object", "properties": {"command": {"type": "string"}},                      "required": ["command"]}},]

Each description says what comes back, not only what the tool does. read_file says the lines are numbered and paged, so the model knows how to read further. run gives two concrete example commands; examples in a description steer the model toward narrow commands, like a single test file, instead of the whole suite. write_file says "the whole new content", which prevents the model from sending a diff that would overwrite the file with half its text.

The toolbox: errors are results

Add the helpers and the dispatcher:

Python
class ToolError(Exception):    """A failure the model should read and recover from."""def clip(text: str, limit: int) -> str:    """Keep the start and, mostly, the end - test summaries and tracebacks live there."""    if len(text) <= limit:        return text    head, tail = text[: limit // 4], text[-(limit * 3 // 4):]    return f"{head}\n... [{len(text) - limit} characters cut] ...\n{tail}"class Toolbox:    def __init__(self, root: Path, cfg, check=None):        self.root, self.cfg = root.resolve(), cfg        self.check = check or (lambda call: None)   # milestone 3 plugs permissions in here        self.handlers = {"read_file": self.read_file, "write_file": self.write_file,                         "run": self.run_command}    def specs(self) -> list[dict]:        return SPECS    def run(self, call: ToolCall) -> ToolResult:        handler = self.handlers.get(call.name)        if handler is None:            return ToolResult(call.id, f"Unknown tool {call.name!r}. Tools: {', '.join(self.handlers)}.", True)        refusal = self.check(call)        if refusal:            return ToolResult(call.id, refusal, True)        try:            return ToolResult(call.id, handler(**call.args))        except ToolError as e:            return ToolResult(call.id, str(e), True)        except TypeError as e:            return ToolResult(call.id, f"Bad arguments for {call.name}: {e}", True)

Toolbox.run never raises for anything the model did. An unknown tool name, a refused call, a ToolError from a handler, or wrong arguments — the TypeError Python raises when a keyword does not match — all come back as a ToolResult with is_error set and a message that says what to do. The naive loop from the anatomy lesson crashed on every one of these.

The check parameter is a hook for the next milestone. It is a function that looks at a call and returns None to allow it, or a refusal message. For now it allows everything. Notice the order: the check runs before the handler, so a refused call never touches the file system.

clip keeps a quarter of its budget for the start of the output and three quarters for the end. That split comes from how test runners and tracebacks are written: the command line and the first failure name are at the top, and the error message, the failing assertion and the summary line are at the bottom. The marker in the middle says how much was cut, so the model knows the output is incomplete.

The three handlers

Add the handlers to Toolbox:

Python
    def _path(self, path: str) -> Path:        full = (self.root / path).resolve()        if not full.is_relative_to(self.root):            raise ToolError(f"{path} is outside the repository. Use paths relative to the repo root.")        return full    def read_file(self, path: str, start: int = 1) -> str:        full = self._path(path)        if not full.is_file():            raise ToolError(f"No file at {path}. Run 'git ls-files' to see what exists.")        lines = full.read_text(errors="replace").splitlines()        start = max(start, 1)        chunk = lines[start - 1 : start - 1 + 400]        body = "\n".join(f"{n:5}  {line}" for n, line in enumerate(chunk, start))        end = start + len(chunk) - 1        return body + (f"\n[lines {start}-{end} of {len(lines)}]" if end < len(lines) else "")    def write_file(self, path: str, content: str) -> str:        full = self._path(path)        full.parent.mkdir(parents=True, exist_ok=True)        full.write_text(content)        return f"Wrote {len(content.splitlines())} lines to {path}."    def run_command(self, command: str) -> str:        try:            proc = subprocess.run(command, shell=True, cwd=self.root, capture_output=True,                                  text=True, timeout=self.cfg.command_timeout)        except subprocess.TimeoutExpired:            raise ToolError(f"Timed out after {self.cfg.command_timeout}s: {command}. "                            "Run something narrower, such as one test file.")        return f"exit code {proc.returncode}\n{clip(proc.stdout + proc.stderr, self.cfg.output_limit)}"

_path resolves the path, following .. and symbolic links, and refuses anything outside the repository root. That stops ../../.ssh/id_rsa, though as the hard-limits lesson said, the run tool can still reach anything the process can; the sandbox is the real boundary.

read_file numbers its lines, so the model can say "line 212" and you can check it, and pages them 400 at a time with a footer that says where it is: [lines 1-400 of 640]. Ledgerly's utils.py is 640 lines; without paging, one read would be most of 10,000 tokens.

run_command treats a non-zero exit code as a normal result, not an error. A failing test is exactly the information the agent asked for; is_error is for when the tool itself could not do its job. A timeout is such a case, and its message suggests the fix.

Test it

Create tests/test_tools.py:

Python
from kite.config import Configfrom kite.model import ToolCallfrom kite.tools import Toolboxdef test_tools_return_results_a_model_can_use(tmp_path):    box = Toolbox(tmp_path, Config())    wrote = box.run(ToolCall("1", "write_file", {"path": "ledgerly/fees.py", "content": "RATE = 2\n"}))    assert wrote.output == "Wrote 1 lines to ledgerly/fees.py."    assert box.run(ToolCall("2", "read_file", {"path": "ledgerly/fees.py"})).output == "    1  RATE = 2"    ran = box.run(ToolCall("3", "run", {"command": "python -c 'print(6 * 7)'"}))    assert ran.output == "exit code 0\n42\n" and not ran.is_errordef test_mistakes_come_back_as_errors_not_crashes(tmp_path):    box = Toolbox(tmp_path, Config())    missing = box.run(ToolCall("1", "read_file", {"path": "ledgerly/nope.py"}))    escape = box.run(ToolCall("2", "write_file", {"path": "../outside.txt", "content": "x"}))    bad_args = box.run(ToolCall("3", "run", {"cmd": "ls"}))    assert missing.is_error and "git ls-files" in missing.output    assert escape.is_error and "outside the repository" in escape.output    assert bad_args.is_error and "Bad arguments" in bad_args.output

Run the tests with your virtual environment active, so that python inside the tool's shell is the right one. You can now run a real read-only task on Ledgerly: build a Toolbox(Path("../ledgerly"), cfg) and call run_session with a question such as "How many test functions does tests/unit/test_fees.py define? Do not change any file." Do not give it write tasks yet. Until the next milestone there are no permissions, and run will execute anything.

Check your understanding

0 of 3 answered

1.A test run exits with code 1 because two tests fail. What does Kite's run tool return?

2.Why does clip keep three quarters of its budget for the end of the output?

3.The model calls write_file with {"file": "tests/test_fees.py", "content": "..."}. What happens?