Course Content
Live Coding Interview Prep
7 sections · 50 lessons
Implement guardrails for tool usage in agents.
What you need to know
The model's tool calls are untrusted input. A prompt injection hidden in a web page, or a plain model mistake, can produce run_sql("DELETE FROM orders"). A line in the system prompt saying "never delete" reduces the chance; only code can make it impossible.
The layers, from coarse to fine:
- Allowlist per role — a support agent can call
lookup_order, notissue_refund. - Rate and spend limits — at most N calls of a tool per session.
- Human confirmation — destructive or irreversible tools wait for a person.
- Argument policy — the arguments themselves are checked: SQL must be a single read-only statement, file paths must stay in the sandbox.
- Audit log — every attempt, allowed or denied, with the reason.
Path traversal is the classic argument attack: reports/../../etc/passwd or simply /etc/passwd. The reliable check is to resolve the path to an absolute one and confirm it is still inside the sandbox directory, not to search for .. in the string.
1import re, time2from collections import defaultdict3from pathlib import Path45SANDBOX = Path("/srv/agent-files").resolve()6WRITE_SQL = re.compile(r"\b(insert|update|delete|drop|alter|create|truncate|grant|merge)\b", re.I)78def validate_args(name: str, args: dict) -> tuple[bool, str]:9 if name == "run_sql":10 sql = str(args.get("query", "")).strip().rstrip(";")11 if ";" in sql:12 return False, "only one statement is allowed"13 if not re.match(r"(?is)^\s*(select|with)\b", sql) or WRITE_SQL.search(sql):14 return False, "only read-only SELECT queries are allowed"15 if name == "read_file":16 target = (SANDBOX / str(args.get("path", ""))).resolve()17 if not target.is_relative_to(SANDBOX):18 return False, "path is outside the sandbox"19 return True, ""2021class ToolGuard:22 """One gate for every tool call: role, limits, confirmation, argument policy, audit."""2324 def __init__(self, allowed: dict[str, set[str]], limits: dict[str, int],25 destructive: frozenset[str] = frozenset()) -> None:26 self.allowed, self.limits, self.destructive = allowed, limits, destructive27 self.counts: defaultdict[str, int] = defaultdict(int)28 self.audit: list[dict] = []2930 def check(self, role: str, name: str, args: dict, confirmed: bool = False) -> tuple[bool, str]:31 if name not in self.allowed.get(role, set()):32 return False, f"tool {name!r} is not permitted for role {role!r}"33 if self.counts[name] >= self.limits.get(name, 10):34 return False, f"call limit reached for {name}"35 if name in self.destructive and not confirmed:36 return False, "human confirmation required"37 return validate_args(name, args)3839 def run(self, role: str, name: str, args: dict, fn, confirmed: bool = False) -> dict:40 ok, reason = self.check(role, name, args, confirmed)41 self.audit.append({"ts": time.time(), "role": role, "tool": name,42 "args": args, "allowed": ok, "reason": reason})43 if not ok:44 return {"ok": False, "error": reason}45 self.counts[name] += 146 try:47 return {"ok": True, "result": fn(**args)}48 except Exception as exc:49 return {"ok": False, "error": f"{type(exc).__name__}: {exc}"}The tricky parts:
startswith("select")is not enough.SELECT 1; DROP TABLE ordersstarts with SELECT. Rejecting a second statement and any write keyword closes that hole; the real backstop is still a database user that only has read permission.resolve()thenis_relative_tocatches../sequences, absolute paths and symlinks alike, because it compares where the path really points, not what the string looks like.- The audit entry is written before the tool runs, so even a call that crashes the process has a record.
- Denials are returned, not raised, so the agent loop shows them to the model, which can choose another approach.
Complexity: role, limit and confirmation checks are O(1) dictionary and set lookups. The SQL regexes are O(length of the query); path resolution is O(path length) plus a filesystem call. The audit log grows O(calls) and should stream to your logging system, not stay in memory.
A real-life example
A support-bot role that may read files and run SQL, and must confirm before refunds:
1guard = ToolGuard(allowed={"support": {"run_sql", "read_file", "issue_refund"}},2 limits={"issue_refund": 2}, destructive=frozenset({"issue_refund"}))3cases = [("run_sql", {"query": "SELECT status FROM orders WHERE id = 42"}, False),4 ("run_sql", {"query": "SELECT 1; DROP TABLE orders"}, False),5 ("run_sql", {"query": "WITH x AS (DELETE FROM orders RETURNING *) SELECT * FROM x"}, False),6 ("read_file", {"path": "reports/march.csv"}, False),7 ("read_file", {"path": "/etc/passwd"}, False),8 ("read_file", {"path": "reports/../../../etc/passwd"}, False),9 ("issue_refund", {"order_id": 42}, False),10 ("issue_refund", {"order_id": 42}, True),11 ("send_email", {"to": "x@y.com"}, False)]12for name, args, confirmed in cases:13 print(name, guard.check("support", name, args, confirmed))| call | decision | reason |
|---|---|---|
| SELECT status … WHERE id = 42 | allow | single read-only statement |
| SELECT 1; DROP TABLE orders | deny | only one statement is allowed |
| WITH x AS (DELETE …) SELECT … | deny | starts with WITH, but contains DELETE |
| read_file reports/march.csv | allow | resolves inside /srv/agent-files |
| read_file /etc/passwd | deny | an absolute path replaces the sandbox root |
| read_file reports/../../../etc/passwd | deny | resolves to /etc/passwd |
| issue_refund, not confirmed | deny | human confirmation required |
| issue_refund, confirmed | allow | confirmed and under the limit of 2 |
| send_email | deny | not in the support role's allowlist |
Note the /etc/passwd row: joining an absolute path onto a Path discards the sandbox root, which is exactly why the check compares the resolved result with the sandbox.
A fintech support agent that can look up transactions but needs an operator's click to reverse one is built on this wrapper.
Follow-up questions to expect
- "Isn't the regex SQL check easy to bypass?" — It raises the bar, but the real control is a database role with only
SELECTgrants on specific views, plus a statement timeout. Code checks are defence in depth, not the only layer. - "How do you stop users clicking 'confirm' on everything?" — Keep the destructive set small, show exactly what will happen ("refund Rs 1,499 to UPI id …"), and put hard limits on amounts that no confirmation can override.
- "How do you test guardrails?" — As a table like the one above, including red-team rows (traversal, stacked SQL, encoded paths), run in CI on every change.