Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your agent's send_email tool accidentally emails 12,000 customers with a test message during a debug session. How do you design agent action safety so destructive actions need approval?
What you need to know
The prompt is not a safety control. Safety comes from what the code allows the agent to do, in which environment, and with whose approval.
Classify every tool
| Class | Examples | Control |
|---|---|---|
| Read-only | Search orders, read a ticket | Run automatically |
| Reversible write | Update a draft, create a ticket, add a tag | Run, log, offer undo |
| Irreversible or high fan-out | Send email, refund, delete, change more than N records | Require explicit approval |
Propose, approve, execute
- Propose — the tool does not send. It returns a pending action: recipient count, a rendered preview, and a signed action token.
- Review — a human sees "Send 'Test message' to 12,000 customers?" with the preview.
- Approve — only an approval with a valid token triggers the executor.
- Execute and audit — the executor sends, and records who approved what and when.
1def send_email(to: list[str], subject: str, body: str, ctx: RunContext) -> dict:2 if ctx.env != "production":3 return sandbox_mail.send(to, subject, body) # debug can never reach real SMTP4 if len(to) > 10 or ctx.tool_policy("send_email") == "approve":5 action = pending_actions.create(tool="send_email", count=len(to),6 preview=render(subject, body), run_id=ctx.run_id)7 return {"status": "pending_approval", "action_id": action.id, "recipients": len(to)}8 return mailer.send(to, subject, body, idempotency_key=f"{ctx.run_id}:{hash(body)}")In LangGraph the same idea is an interrupt before the tool node; in a hand-written loop it is a pending-actions table and a resume endpoint.
Controls that would have stopped this incident
- Environment isolation. Debug and staging sessions get a mock email tool bound to a sandbox domain and no production credentials. The test run should have been incapable of reaching real customers. This is a configuration boundary, not a matter of discipline.
- Quantity thresholds. More than 10 recipients needs approval, whoever calls it.
- Rate limits per agent, per tool, per hour, enforced in the executor.
- A kill switch. One flag disables all write tools instantly. New tools default to dry-run mode.
Audit and watch
Log every action: run id, tool, arguments, approver, result. Track approval-queue latency, the share of actions auto-approved, and the rejection rate. A rising rejection rate means the agent's judgement is drifting.
A real-life example
Scenario (illustrative numbers). A subscription-box company's marketing agent can draft and send campaign emails. An engineer debugging a prompt locally runs the agent against production config, and a test email reaches 12,000 customers before anyone notices. Unsubscribes jump by 900 that day.
The fix has four parts. Local and staging runs now load a sandbox mail tool with no production credentials. send_email returns a pending action above 10 recipients. The executor enforces 2,000 emails per hour per agent. And a kill switch in the admin panel disables all write tools. In the next quarter, 340 campaign sends go through approval with a median wait of 6 minutes, and 12 are rejected, 3 of them for wrong audience segments.
Follow-up questions to expect
- "Won't approvals slow everything down?" — Only for the risky class. Most actions are reads or reversible writes; approvals apply to a small share, and a good preview makes review quick.
- "Can the model approve its own action?" — No. The approval must come from a person, or a separate policy service, through a channel the model cannot call.
- "What about prompt injection asking the agent to email everyone?" — The same gates apply; injection can create a proposal, but not an execution.