Course Content
Harness Engineering: Making Coding Agents Dependable
5 sections · 23 lessons
Why every session starts from zero
On Friday evening, an agent session on Ledgerly's feature F2 — "reminder emails show the late fee and the new total" — ran out of turns at 60. It had found the right template, added the fee line, and spent its last 20 turns fighting a test that would not capture the email. On Monday morning a new session started on the same feature.
The Monday session re-read the same 12 files. It tried the same approach that had failed on Friday: patching smtplib.SMTP in the test. It found Friday's half-finished template change, did not recognise it as its own earlier work, and "cleaned it up" by reverting it. By turn 22 it was roughly where Friday's session had been at turn 30.
Nothing was wrong with the model on Monday. It simply knew nothing about Friday. This lesson explains why that is the normal state of affairs, why it is actually useful, and what the harness must write down to make it work.
What survives and what does not
A model has no memory between calls, and a session is only a message list. When the session ends, the list is gone unless something saves it. Here is what the Monday session could and could not know:
| Survives into the next session | Lost when the session ends |
|---|---|
| Files on disk, including half-finished edits | Why each edit was made |
| Git history and commit messages | Approaches that were tried and failed |
| Anything written to a progress or notes file | Decisions made along the way ("fee applies to the unpaid balance") |
| The instruction file | Which files were already read and understood |
| Logs, if the harness kept them | How close the work was to finished |
The right column is the expensive part. A failed approach costs 10 turns to discover and nothing to write down. If nobody writes it down, the next session pays the 10 turns again.
Three ways to continue work
When a session stops before the work is done, you have three options.
Resume the conversation
- Reload the full message list and keep going
- Nothing is lost, including every dead end
- The next turn reads 180,000 tokens: about 0.90 USD uncached, every turn
- Drift and stale assumptions come back with it
Compact it
- Replace old messages with a model-written summary
- Much smaller context
- The summary is written by the session that got stuck
- What it drops, you cannot see or check
Start fresh with a handoff
- New session, written briefing, clean context
- Starts at about 6,000 tokens
- Only as good as what was written down
- Needs the harness to write the handoff reliably
Resuming makes sense for short interruptions: a session paused for a human approval, a network error mid-run. Compaction is useful inside a long session when the context grows large, and many tools do it automatically. But for carrying work across a real break — a failed session, the next day, the next feature — starting fresh is usually best.
Why a fresh start is a feature
It is tempting to see "starts from zero" as a weakness to engineer around. For feature work it is closer to a strength.
- The task is in focus. A fresh session's context holds the task, the briefing and the instruction file. Nothing from Friday's frustrating test fight is pulling at its attention. The drift mechanism from the first section is reset.
- Mistakes do not carry over. A wrong assumption from turn 8 of Friday's session is not in Monday's context unless someone wrote it down as fact.
- Runs are comparable. Every session starts from the same kind of input: repo state plus briefing. You can compare two sessions, replay one, or run two on different features at the same time.
- The handoff forces clarity. Writing down "patching smtplib does not work, patch ledgerly.mail.send" is a small act of debugging. It turns a vague struggle into one sentence of knowledge.
The cost is the re-reading. A fresh Ledgerly session spends 3 to 5 turns and about 15,000 tokens reading the files it needs. That is real, but it is far less than the 0.90 USD per turn of dragging a 180,000-token history through the rest of a resumed session.
What the handoff must contain
A useful handoff answers four questions. What is the overall goal? What is already done and verified? What exactly is this session's task and how will it be checked? What did earlier attempts learn?
Kite builds this briefing from its progress file at the start of every session. Here is Monday's briefing for F2:
Project: Ledgerly: late fees on overdue invoices, shown everywhere a customer sees money.Already done and verified by the harness (do not redo):- F1 2% late fee on invoices unpaid more than 30 days after the due dateYour task this session: F2 - Reminder emails show the late fee and the new totalIt is accepted when this passes: python -m pytest -q tests/unit/test_reminders.py, and so do the standard checks.Notes from earlier attempts: Template change in templates/email/reminder.txt is doneand correct. Patching smtplib.SMTP in tests does not work, because reminders.py sendsthrough ledgerly.mail.send; patch that instead. Take the fee from fees.apply_late_fee;do not recompute it in the template.The notes paragraph is the most valuable text in it. It comes from the previous session's final message, which Kite saves automatically. That is why Kite's system prompt asks the agent to finish with a short summary of what it changed. When a session fails, the same summary becomes the next session's head start. A good system prompt for a harness asks for this handoff explicitly: what was done, what failed and why, and what to try next.
Notice where the handoff does not live: in a chat window, in the agent's own scratch file, or only in the model's context. It lives in a file the harness owns, next to the code, and under version control.
When a long session is the right call
Starting fresh is the default for feature work, not a law. Some tasks need one continuous line of thought: chasing a race condition where each experiment depends on the last, or a migration where the agent must hold a whole plan in mind. For those, give the session a bigger turn budget, let compaction handle the growth, and watch the log closely. Just do not let "some tasks need long sessions" become the reason every task gets one.
Check your understanding
0 of 3 answered
1.A session on F3 fails after 60 turns. You want to continue tomorrow. Which approach is usually best for this kind of feature work?
2.Which piece of information is most likely to be lost between sessions if nobody writes it down?
3.Why does Kite store the handoff in a harness-owned file in the repository rather than in the agent's own notes file?