Harness Engineering: Making Coding Agents Dependable

Set up the project before the first feature


The first Kite session on Ledgerly was asked to build F1, the late fee. It spent 19 of its 31 turns not on the late fee but on getting the tests to run at all. A dependency was missing from the dev requirements. The SQLite test database path pointed at a folder that did not exist on a fresh clone. The app refused to start without a SECRET_KEY environment variable. The agent fixed all three, then built the late fee in a hurry with 12 turns left.

The result was one commit that mixed three setup fixes with a feature. The late fee could not be reviewed on its own. The setup fixes could not be reverted without losing the fee. And the next feature session had to discover the same three problems, because nothing had recorded that they were fixed or how.

Setup and features are different kinds of work, with different goals and different checks. This lesson gives setup its own session.

What must exist before feature workGreen baseline from a fresh clonescripts/check.sh — the one checkAGENTS.md, under 60 lineskite-progress.json with verify commandsA human has reviewed the feature plan
Doing setup in its own session moved the first surviving edit from turn 19 to turn 5 and gave every feature a clean commit.

What initialisation produces

An initialisation session has one goal: make the repository ready for feature sessions. On Ledgerly, "ready" means five artifacts exist and work:

ArtifactWhy feature sessions need it
A green baseline: bash scripts/check.sh passes on a clean checkoutThe gate cannot tell new failures from old ones if the baseline is red
scripts/check.sh: one command for the full checkThe gate, CI and humans all run the same thing
A documented fast test commandThe inner loop stays around 9 seconds, not 41
AGENTS.md: a short root instruction fileEvery session starts with the commands and the hard rules
kite-progress.json: the goal and a feature list with verify commandsFeature sessions know what to do and how they will be judged

The first row is the one teams skip, and it matters most. If the suite already fails before any agent touches the code — say, because of the flaky PDF test — then every feature session's gate fails for a reason that has nothing to do with the feature. Initialisation is where the flaky test gets its marker and its quarantine, once, by a session whose job is exactly that.

The initialisation task

The task text for an initialisation session is different from a feature task. It asks for preparation, not features, and it forbids changing behaviour:

Text
Prepare this repository for feature work by coding agents. Do not add featuresor change application behaviour.1. Make the test suite runnable from a fresh clone. Fix setup problems only   (missing dev dependencies, paths, required environment variables with safe   test defaults). Record each fix in your final message.2. Create scripts/check.sh that runs: ruff check ledgerly tests, then   python -m pytest -q -x -m "not flaky", then flask --app ledgerly.app routes.3. If a test fails intermittently, mark it @pytest.mark.flaky with a comment   explaining why. Register the marker in pyproject.toml. Do not delete it.4. Write AGENTS.md: under 60 lines, with the commands, the hard rules and where   to look. Do not include information the code already shows.5. Write kite-progress.json for this goal: "Late fees on overdue invoices, shown   everywhere a customer sees money." Split it into features that one session   can finish, each with a verify command naming a test file that will prove it.

Every step names its output. Step 1 says to fix setup problems only, and to report them, so they end up in the commit message and not buried in a feature. Step 5 asks the agent to propose the plan, which a human will review.

The initialisation gate

Initialisation needs its own gate, because "done" means something different here. It is not "the feature works". It is "the repository is ready":

Bash
#!/usr/bin/env bash# scripts/init-gate.sh: is this repository ready for feature sessions?set -euo pipefailtest -f AGENTS.md && [ "$(wc -l < AGENTS.md)" -le 60 ]bash scripts/check.sh                      # the baseline must be greenpython - <<'EOF'import jsondata = json.load(open("kite-progress.json"))assert data["goal"] and data["features"], "goal and features are required"for f in data["features"]:    assert f["status"] == "todo", f"{f['id']} should start as todo"    assert f["verify"].startswith("python -m pytest"), f"{f['id']} needs a pytest verify command"print(f"ready: {len(data['features'])} features")EOF

It checks the three things that feature sessions depend on: the instruction file exists and is short, the baseline is green, and every feature has a runnable verification command. Run it twice from a fresh clone to be sure the baseline is not green by luck — with a flaky test in the suite, one green run proves little.

Kite, as you build it in this course, has no initialisation mode: it expects kite-progress.json to exist, and its permission layer protects that file. Until you add one — it is the exercise at the end of milestone 5 — run initialisation with the agent tool you already use, using the task above, and run the gate script yourself before accepting the result.

A human reviews the plan

The most important output of initialisation is the feature list, and it is a plan, not code. A human must read it before any feature session starts. On Ledgerly, the review took 15 minutes and changed three things:

  • It removed one feature, "Refactor fee calculations into a service class", which was not part of the goal.
  • It rewrote two verify commands that pointed at test files covering far more than the feature.
  • It split "Late fee shown in PDF and email" into two features, because each needed a different kind of test.

This review is cheap and high-leverage. Every later session inherits the plan. A bad feature boundary here costs failed sessions later; a missing verify command makes a feature impossible to gate.

When to skip it

Not every repository needs a full initialisation session. If the tests already run from a clean clone, check.sh exists and an instruction file is in place, a single short session — or a human spending 20 minutes — can write the progress file. Initialisation pays back when the setup is broken or undocumented, or when many feature sessions will follow. On Ledgerly it cost one session, 26 turns and about 3 USD with caching.

Check your understanding

0 of 3 answered

1.Why must the baseline be green before feature sessions start?

2.The initialisation session proposes the feature "Improve the reminder system". What should the human reviewer do?

3.Why does the initialisation gate run check.sh as one of its checks, and why run the gate twice from a fresh clone?