Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
The repository is the agent's whole world
A new engineer joining the Ledgerly team gets a week of help. Someone tells them that utils.py is scary and should not be touched. Someone shows them the command that runs only the fast tests. In a Slack thread they learn that the PDF test is flaky and that migrations are never edited after they ship. Most of what makes them effective never appears in the code.
A coding agent gets none of that. It gets the files on disk, the output of the commands it runs, and whatever you put in its instructions. If a fact is not in one of those three places, then for the agent it does not exist. This lesson is about taking that seriously.
What the agent can and cannot see
The agent can see
- Every file in the working tree it chooses to read
- Git history, if it thinks to run
git log - Output of commands it runs: tests, linters, grep
- Its instruction file and the task text
The agent cannot see
- Slack threads, meetings, tickets, design reviews
- Why the code is the way it is, unless written down
- Production data, dashboards, the staging database
- Which tests are flaky, slow or secretly important
The right column is where most "mysterious" agent decisions come from. The agent that edited migrations/0019_add_gstin.py in place did not know that migration had already run in production. There was no comment in the file and no rule in the instructions. To the agent, that migration was just an ordinary Python file with a bug in it.
The fix is not to write everything down. It is to write down the handful of facts whose absence causes damage, in the place the agent will find them at the moment it needs them. A comment at the top of the migrations folder's README.md reaches every agent that lists that folder.
Navigability: finding things costs turns
An agent does not start with a map of the codebase. It builds one by running ls, grep and reading files, and every step costs a turn and tokens. Discovery cost is a useful measure: how many turns and tokens pass before the agent's first edit that survives into the final diff.
On Ledgerly's late-fee task, discovery cost was 14 turns and 92,000 tokens. The agent spent most of them in utils.py, because fee calculations were spread across do_calc(), fx2() and apply_adj() — names that tell you nothing. It read all 640 lines twice.
Three properties of a repository drive discovery cost:
- Names that say what things do.
apply_late_fee()is found by one grep.fx2()is found by reading everything. - Things in the place you would look. Fee logic in
ledgerly/fees.py, not scattered through a utilities file. - A short map. Ten lines saying which folder holds what save ten turns of
ls.
A map the harness can generate
You can hand the agent a map rather than make it build one. Here is a small script that prints every Python module with its line count and top-level functions and classes. A harness can run it once and put the output in the first message.
1import ast2import sys3from pathlib import Path456def repo_map(root: Path, max_names: int = 8) -> str:7 """One line per module: path, line count, and its top-level names."""8 lines = []9 for path in sorted(root.rglob("*.py")):10 if any(part.startswith(".") or part == "migrations" for part in path.parts):11 continue12 source = path.read_text(errors="replace")13 try:14 tree = ast.parse(source)15 except SyntaxError:16 continue17 names = [node.name for node in tree.body18 if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef, ast.ClassDef))]19 shown = ", ".join(names[:max_names]) + (" ..." if len(names) > max_names else "")20 lines.append(f"{path.relative_to(root)} ({len(source.splitlines())} lines): {shown}")21 return "\n".join(lines)222324if __name__ == "__main__":25 print(repo_map(Path(sys.argv[1])))On Ledgerly this prints 33 lines, about 1,900 tokens. A line like ledgerly/utils.py (640 lines): round_money, parse_date, fmt_inr, fx2, do_calc, apply_adj, legacy_tax_code, _norm ... tells the agent immediately that fee-like logic lives there, and that it is a big file worth searching rather than reading.
The trade-off: a map is a snapshot. On a 40-file repository it is cheap and always helpful. On a 4,000-file monorepo it would be 200,000 tokens, which is worse than no map. There you would map only the package the task touches, or rely on good names and a short written guide instead.
Make the repository easy to work in
The best improvements for agents are the same ones that help a new engineer. On Ledgerly, five changes cut discovery cost on the late-fee task from 14 turns to 4:
- One command for the fast tests —
pytest tests/unit -qruns in 9 seconds and is written in the instruction file, so the agent stops running the 41-second suite after every edit. - Mark the flaky test —
@pytest.mark.flakyon the PDF test, with a comment saying why, and the standard check excludes it. - Move fee logic to one module —
ledgerly/fees.pywithapply_late_fee()andgst_breakup();utils.pyre-exports the old names so nothing else breaks. - A README in
migrations/— "Never edit a migration after it is merged. Add a new one withalembic revision." - Type hints on money functions —
def round_money(amount: Decimal) -> Decimaltells the agent the contract without reading the body.
None of these changes is about AI. Each is ordinary engineering that the team had postponed because humans could work around the problems. Agents cannot work around them, so they show you exactly where your repository's hidden costs are.
Check your understanding
0 of 3 answered
1.An agent keeps running python -m pytest (41 seconds) after every small edit on Ledgerly. What is the most effective fix?
2.What does "discovery cost" measure?
3.Why is a generated repository map a good idea for Ledgerly but a bad one for a 4,000-file monorepo?