Harness Engineering: Making Coding Agents Dependable

Scope control: do it, flag it, or queue it


Every real task uncovers other work. While adding the late fee to Ledgerly, an agent will read utils.py and notice that round_money() uses float arithmetic. It will see a typo in an old migration. It will find that parse_date() accepts six date formats and handles two of them wrongly. All of these are real problems. None of them is the task.

A human engineer has instincts about this. They fix a typo in a line they are already touching. They mention the float bug in the pull request. They do not rewrite utils.py on a Tuesday afternoon without asking. An agent has no such instincts unless you give it rules, and without rules it tends to fix whatever is in front of it. The result is the 540-line diff that nobody wants to review.

This lesson gives the agent three choices for discovered work, and gives the harness a way to make the third choice safe.

Deciding what to do with discovered workWork found mid-taskTask needsit: executeTask does not need itTask iswrong: flag itUnrelated: queue it
Queued items become proposals only a human can promote, so the agent can suggest work but never assign it to itself.

Three choices

  1. Do it — the work is needed to finish the task. The late fee needs a days_overdue() helper; write it. It is in scope even though the task did not mention it.
  2. Flag it — the task as written is wrong, ambiguous or impossible. The task says "2% late fee", but it does not say whether the fee is charged once or every month. Stop and say so. Do not guess.
  3. Queue it — the work is real but not needed for this task. The float bug in round_money() goes into a list for a human to review. Do not fix it now.

The test for doing it is simple: would the task be incomplete without this change? If yes, do it. The test for flagging it: would a reasonable reviewer disagree with my reading of the task? If yes, stop and flag it. Everything else is queued.

Flagging deserves care in unattended runs, where nobody can answer a question. The agent should not guess and continue. Kite handles this by letting the agent finish the session with a clear explanation instead of an implementation. The session ends as "not done", the note goes into the progress file, and a human clarifies the task before the next attempt. One wasted session is much cheaper than a confident implementation of the wrong rule.

Why queue instead of fix?

Fixing things you notice feels responsible. For an agent it is usually a mistake, for three reasons.

  • Review cost. A reviewer who asked for a late fee must now also check a change to rounding that affects every invoice. The float fix touches 17 importers of utils.py.
  • Hidden risk. The rounding fix changes totals by a paisa on some invoices. That may be correct and still need a customer announcement. The agent cannot know that.
  • Attribution. If the release breaks, "late fee" is the change people will look at. A rounding change hidden in it makes debugging slower.

Queue is not "ignore". A queued item goes to a human who decides whether it becomes its own task, with its own tests and its own review.

Making the queue real

A rule that says "report unrelated problems" is only as good as what happens to the report. If queued items end up in a chat transcript nobody reads, the agent learns nothing and neither do you. Kite asks the agent to put each item on its own line starting with QUEUE: in its final message, and the harness copies those lines into the progress file as proposed features:

Python
for line in outcome.summary.splitlines():          # scope control: queue it, never do it    if line.startswith("QUEUE:"):        data["features"].append({"id": f"Q{len(data['features']) + 1}",                                 "title": line[6:].strip(), "status": "proposed",                                 "verify": "", "notes": f"found during {feature['id']}"})

This comes from Kite's progress.py, which you build in milestone 5. The status proposed matters. Kite only picks up features with status todo or in_progress, so a proposed item waits until a human reads it, writes a verification command for it, and promotes it. The agent can suggest work. It cannot assign work to itself.

Diff size as a scope signal

You can often see a scope problem from the size and shape of the diff before reading a line of it. The late-fee task on Ledgerly should touch about four files: fees.py, one service file, one test file, maybe one new migration. The unguarded run touched 11 files and 540 lines.

A cheap check compares changed paths with the paths you expected:

Python
import subprocessfrom pathlib import Pathdef out_of_scope(root: Path, expected: list[str]) -> list[str]:    """Changed files that are not under any expected path prefix."""    changed = subprocess.run(["git", "diff", "--name-only", "HEAD"], cwd=root,                             capture_output=True, text=True, check=True).stdout.split()    return [path for path in changed if not path.startswith(tuple(expected))]print(out_of_scope(Path("../ledgerly"), ["ledgerly/fees.py", "ledgerly/invoices/", "tests/"]))

On the unguarded run this prints ['ledgerly/utils.py', 'migrations/0019_add_gstin.py', 'ledgerly/pdf.py', ...]. Each path is a question for the reviewer, or a reason for the harness to fail the session.

Do not make this a hard gate by default. Expected paths are a guess, and legitimate work sometimes lands somewhere you did not predict: the late fee might reasonably need a change in reminders.py. Use it as a warning in the log and in the review summary. Turn it into a hard failure only for paths that should never change during feature work.

When wide scope is the task

Scope rules are for feature and bug-fix tasks. Some tasks are about breadth: "rename fx2 to convert_currency everywhere", "move all money math to Decimal". For those, the scope is the whole codebase by design, and a queue rule would just get in the way. Say so in the task text, and lean harder on the gate and on git checkpoints instead, because a wide change has more places to go wrong.

Check your understanding

0 of 3 answered

1.While adding a CSV export, the agent finds that the export needs invoices sorted by date, and Invoice has no index on issued_at, so the query takes 4 seconds on production-sized data. What is the right choice?

2.The task says "Add a late fee of 2%." During the run it becomes clear the task does not say whether the fee is one-time or monthly. Kite is running unattended. What should the agent do?

3.Why does Kite give queued items the status proposed rather than todo?