Harness Engineering: Making Coding Agents Dependable

Hard limits: permissions, deny lists and sandboxes


Two real incidents from Ledgerly's first month with agents. In the first, an agent working on reminder emails could not get the test SMTP settings to load. It read .env, found SMTP_PASSWORD, and copied the real value into a test fixture "so the tests are self-contained". The fixture was committed to a branch and pushed before anyone noticed. The password had to be rotated.

In the second, an agent wanted "a clean state" after a messy attempt and ran git reset --hard HEAD. That also wiped two hours of a human engineer's uncommitted work in the same checkout. The agent had done exactly what the command does. Nobody had told it that it was sharing the working tree with a person.

Neither agent was malicious or broken. Each took a reasonable-looking step with a large blast radius. Instructions help with this, but for actions like these, "usually follows the rule" is not enough. This lesson is about limits the harness enforces, and about being honest regarding how strong those limits are.

Limits, from the wall upSandbox:worktree, fake secretsDeny: push,reset --hard, rm -rfAsk:migrations, pip, unknownAllow: read,tests, lint, difftopbottomAny code the agent writes and runs through pytest bypasses the rules above the sandbox.
Permission tiers catch a cooperative agent's mistakes; only the sandbox at the bottom limits what a bad session can reach.

Three tiers

Every action the agent can take falls into one of three tiers:

TierMeaningLedgerly examples
AllowRuns without askingRead any file except secrets; write inside ledgerly/ and tests/; run tests, lint, git status, git diff, grep
AskA human approves first; refused when unattendedWrite under migrations/, .github/, pyproject.toml; pip install; any command not on the safe list
DenyNever runs, even if a human would approve in the momentgit push, git reset --hard, git clean, rm -rf, sudo, piping a download into a shell

The ask tier is where the design work is. Too little in it and dangerous actions slip through. Too much, and a human approving dozens of prompts per session starts clicking "yes" without reading, which is worse than no prompt. On Ledgerly, a well-tuned policy asks about 2 to 4 times per feature session. More than 10 means the allow list is too tight.

The deny tier is for actions whose damage is hard to undo and which the agent never legitimately needs. git push is the classic case: Kite commits locally, and a human pushes after review. Deny rules are not about distrust. They make sure a tired human cannot approve a disaster at 7 p.m.

What the rules look like

Here are Kite's command rules from permissions.py, which you build in milestone 3:

Python
DENY = [r"\brm\s+-\w*[rf]", r"\bgit\s+(push|reset\s+--hard|clean)\b", r"\bsudo\b",        r"\bcurl\b.*\|\s*(ba|z)?sh\b", r"\bchmod\s+-R\b", r"\bdrop\s+(table|database)\b",        r"\.env\b"]                          # secrets: never cat, grep or copy .envSAFE = re.compile(r"(python -m pytest|pytest|ruff|ls|cat|head|tail|wc|grep|rg|"                  r"git (status|diff|log|show|ls-files))\b")TRICKS = re.compile(r"[;&|`$<>]")          # chaining, pipes, substitution, redirectsPROTECTED = ("migrations/", ".env", ".github/", "pyproject.toml", "kite-progress.json")

A command is denied if any deny pattern matches anywhere in it, which is how cat .env is stopped; the read_file tool refuses .env separately. It is allowed only if it starts with a safe command and contains no shell tricks. Everything else is asked. The tricks rule matters: without it, pytest -q; rm -rf build starts with a safe command and would sail through. With it, any chaining, piping or redirection drops the command to "ask".

The same tiers exist in the tools you already use. In Claude Code, a project's .claude/settings.json can hold permission rules like this:

JSON
{  "permissions": {    "allow": ["Bash(python -m pytest:*)", "Bash(ruff check:*)", "Bash(git diff:*)"],    "deny": ["Bash(git push:*)", "Read(./.env)", "Read(./.env.*)"]  }}

Check your tool's documentation for the exact rule syntax in your version. The design question — what goes in each tier — is the same everywhere.

A deny list is a speed bump, not a wall

Be honest about what pattern rules can do. They catch mistakes by a cooperative agent: the agent that thinks git reset --hard is a harmless cleanup. They do not stop an agent that has been steered into doing harm — for example by instructions hidden in a file it read, which is called prompt injection.

The reason is simple. Once an agent can write files and run tests, it can run any code. It can write a test file that deletes a directory with shutil.rmtree, then run python -m pytest, which is on the safe list. It can write a script that reads .env and prints it. No regular expression over command strings closes these paths, because the dangerous part is inside code the agent wrote.

Permission rules

  • Catch honest mistakes before they run
  • Cheap, fast, easy to read and test
  • Give the agent a useful refusal message
  • Bypassed by any code the agent writes and runs

A sandbox

  • Limits what any code can reach
  • No real secrets, no production network, no shared checkout
  • Makes the worst case a thrown-away container
  • Costs setup work and some convenience

You want both. The rules make normal sessions smooth and safe. The sandbox makes the rare bad session harmless.

A sandbox that makes the worst case boring

For Ledgerly, the sandbox has four properties: a fresh checkout, no real secrets, no access to anything you would miss, and a user without special rights.

Bash
# A clean worktree: it has the code but not your uncommitted work or your .envgit worktree add ../ledgerly-agent -b agent/late-fee# A throwaway container: fake settings, non-root user, repo mounted, nothing elsedocker run --rm -it \  --user 1000:1000 \  -v "$(realpath ../ledgerly-agent)":/work -w /work \  -e ANTHROPIC_API_KEY \  -e DATABASE_URL=sqlite:///dev.db -e SMTP_PASSWORD=not-a-real-password \  ledgerly-agent-image python -m kite /work --unattended

The worktree solves the git reset --hard incident: the agent has its own checkout, so it cannot touch a human's uncommitted work. The fake settings solve the .env incident: there is no real password to leak. The container limits the damage of anything else to one disposable directory. The API key is the one real secret inside; give the agent a key with a spending limit.

What this setup does not do is restrict the network. The container can still reach the internet, including anything an injected instruction might point it at. Tighter setups route traffic through a proxy that allows only the model API and your package index, or run the harness outside the container and only the tools inside it. Decide how far to go from what is reachable: a laptop with production credentials in the shell deserves far more care than a CI runner with none.

Check your understanding

0 of 3 answered

1.Kite allows python -m pytest without asking. An agent writes tests/test_cleanup.py that deletes the ledgerly/ folder, then runs python -m pytest. What stops it?

2.A human approving Kite's prompts is asked 25 times per session. What is the most likely problem?

3.Why run the agent in a separate git worktree instead of your normal checkout?