Live Coding Interview Prep

Course Content

Live Coding Interview Prep

7 sections · 50 lessons

Implement guardrails for tool usage in agents.


One gate every tool call passesRole allowlistPer-sessioncall limitHumanconfirmationif destructiveArgumentpolicy:SQL, pathsAudit, then run'SELECT 1; DROP TABLE' and '/etc/passwd' both stop at the argument policy.
The prompt can ask the model to behave; only this wrapper can make a dangerous call impossible.

What you need to know

The model's tool calls are untrusted input. A prompt injection hidden in a web page, or a plain model mistake, can produce run_sql("DELETE FROM orders"). A line in the system prompt saying "never delete" reduces the chance; only code can make it impossible.

The layers, from coarse to fine:

  1. Allowlist per role — a support agent can call lookup_order, not issue_refund.
  2. Rate and spend limits — at most N calls of a tool per session.
  3. Human confirmation — destructive or irreversible tools wait for a person.
  4. Argument policy — the arguments themselves are checked: SQL must be a single read-only statement, file paths must stay in the sandbox.
  5. Audit log — every attempt, allowed or denied, with the reason.

Path traversal is the classic argument attack: reports/../../etc/passwd or simply /etc/passwd. The reliable check is to resolve the path to an absolute one and confirm it is still inside the sandbox directory, not to search for .. in the string.

Python
import re, timefrom collections import defaultdictfrom pathlib import PathSANDBOX = Path("/srv/agent-files").resolve()WRITE_SQL = re.compile(r"\b(insert|update|delete|drop|alter|create|truncate|grant|merge)\b", re.I)def validate_args(name: str, args: dict) -> tuple[bool, str]:    if name == "run_sql":        sql = str(args.get("query", "")).strip().rstrip(";")        if ";" in sql:            return False, "only one statement is allowed"        if not re.match(r"(?is)^\s*(select|with)\b", sql) or WRITE_SQL.search(sql):            return False, "only read-only SELECT queries are allowed"    if name == "read_file":        target = (SANDBOX / str(args.get("path", ""))).resolve()        if not target.is_relative_to(SANDBOX):            return False, "path is outside the sandbox"    return True, ""class ToolGuard:    """One gate for every tool call: role, limits, confirmation, argument policy, audit."""    def __init__(self, allowed: dict[str, set[str]], limits: dict[str, int],                 destructive: frozenset[str] = frozenset()) -> None:        self.allowed, self.limits, self.destructive = allowed, limits, destructive        self.counts: defaultdict[str, int] = defaultdict(int)        self.audit: list[dict] = []    def check(self, role: str, name: str, args: dict, confirmed: bool = False) -> tuple[bool, str]:        if name not in self.allowed.get(role, set()):            return False, f"tool {name!r} is not permitted for role {role!r}"        if self.counts[name] >= self.limits.get(name, 10):            return False, f"call limit reached for {name}"        if name in self.destructive and not confirmed:            return False, "human confirmation required"        return validate_args(name, args)    def run(self, role: str, name: str, args: dict, fn, confirmed: bool = False) -> dict:        ok, reason = self.check(role, name, args, confirmed)        self.audit.append({"ts": time.time(), "role": role, "tool": name,                           "args": args, "allowed": ok, "reason": reason})        if not ok:            return {"ok": False, "error": reason}        self.counts[name] += 1        try:            return {"ok": True, "result": fn(**args)}        except Exception as exc:            return {"ok": False, "error": f"{type(exc).__name__}: {exc}"}

The tricky parts:

  • startswith("select") is not enough. SELECT 1; DROP TABLE orders starts with SELECT. Rejecting a second statement and any write keyword closes that hole; the real backstop is still a database user that only has read permission.
  • resolve() then is_relative_to catches ../ sequences, absolute paths and symlinks alike, because it compares where the path really points, not what the string looks like.
  • The audit entry is written before the tool runs, so even a call that crashes the process has a record.
  • Denials are returned, not raised, so the agent loop shows them to the model, which can choose another approach.

Complexity: role, limit and confirmation checks are O(1) dictionary and set lookups. The SQL regexes are O(length of the query); path resolution is O(path length) plus a filesystem call. The audit log grows O(calls) and should stream to your logging system, not stay in memory.

A real-life example

A support-bot role that may read files and run SQL, and must confirm before refunds:

Python
guard = ToolGuard(allowed={"support": {"run_sql", "read_file", "issue_refund"}},                  limits={"issue_refund": 2}, destructive=frozenset({"issue_refund"}))cases = [("run_sql", {"query": "SELECT status FROM orders WHERE id = 42"}, False),         ("run_sql", {"query": "SELECT 1; DROP TABLE orders"}, False),         ("run_sql", {"query": "WITH x AS (DELETE FROM orders RETURNING *) SELECT * FROM x"}, False),         ("read_file", {"path": "reports/march.csv"}, False),         ("read_file", {"path": "/etc/passwd"}, False),         ("read_file", {"path": "reports/../../../etc/passwd"}, False),         ("issue_refund", {"order_id": 42}, False),         ("issue_refund", {"order_id": 42}, True),         ("send_email", {"to": "x@y.com"}, False)]for name, args, confirmed in cases:    print(name, guard.check("support", name, args, confirmed))
calldecisionreason
SELECT status … WHERE id = 42allowsingle read-only statement
SELECT 1; DROP TABLE ordersdenyonly one statement is allowed
WITH x AS (DELETE …) SELECT …denystarts with WITH, but contains DELETE
read_file reports/march.csvallowresolves inside /srv/agent-files
read_file /etc/passwddenyan absolute path replaces the sandbox root
read_file reports/../../../etc/passwddenyresolves to /etc/passwd
issue_refund, not confirmeddenyhuman confirmation required
issue_refund, confirmedallowconfirmed and under the limit of 2
send_emaildenynot in the support role's allowlist

Note the /etc/passwd row: joining an absolute path onto a Path discards the sandbox root, which is exactly why the check compares the resolved result with the sandbox.

A fintech support agent that can look up transactions but needs an operator's click to reverse one is built on this wrapper.

Follow-up questions to expect

  • "Isn't the regex SQL check easy to bypass?" — It raises the bar, but the real control is a database role with only SELECT grants on specific views, plus a statement timeout. Code checks are defence in depth, not the only layer.
  • "How do you stop users clicking 'confirm' on everything?" — Keep the destructive set small, show exactly what will happen ("refund Rs 1,499 to UPI id …"), and put hard limits on amounts that no confirmation can override.
  • "How do you test guardrails?" — As a table like the one above, including red-team rows (traversal, stacked SQL, encoded paths), run in CI on every change.