Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your agent's send_email tool accidentally emails 12,000 customers with a test message during a debug session. How do you design agent action safety so destructive actions need approval?


The tool's blast radius decides the controlRead-only: run automaticallyReversible write: run, log, offer undoOver 10 recipients: pending approvalDebug session: sandbox mail, no SMTP
The test email could only reach customers because a debug run held production credentials; the fix is a boundary, not a reminder.

What you need to know

The prompt is not a safety control. Safety comes from what the code allows the agent to do, in which environment, and with whose approval.

Classify every tool

ClassExamplesControl
Read-onlySearch orders, read a ticketRun automatically
Reversible writeUpdate a draft, create a ticket, add a tagRun, log, offer undo
Irreversible or high fan-outSend email, refund, delete, change more than N recordsRequire explicit approval

Propose, approve, execute

  1. Propose — the tool does not send. It returns a pending action: recipient count, a rendered preview, and a signed action token.
  2. Review — a human sees "Send 'Test message' to 12,000 customers?" with the preview.
  3. Approve — only an approval with a valid token triggers the executor.
  4. Execute and audit — the executor sends, and records who approved what and when.
Python
def send_email(to: list[str], subject: str, body: str, ctx: RunContext) -> dict:    if ctx.env != "production":        return sandbox_mail.send(to, subject, body)            # debug can never reach real SMTP    if len(to) > 10 or ctx.tool_policy("send_email") == "approve":        action = pending_actions.create(tool="send_email", count=len(to),                                        preview=render(subject, body), run_id=ctx.run_id)        return {"status": "pending_approval", "action_id": action.id, "recipients": len(to)}    return mailer.send(to, subject, body, idempotency_key=f"{ctx.run_id}:{hash(body)}")

In LangGraph the same idea is an interrupt before the tool node; in a hand-written loop it is a pending-actions table and a resume endpoint.

Controls that would have stopped this incident

  • Environment isolation. Debug and staging sessions get a mock email tool bound to a sandbox domain and no production credentials. The test run should have been incapable of reaching real customers. This is a configuration boundary, not a matter of discipline.
  • Quantity thresholds. More than 10 recipients needs approval, whoever calls it.
  • Rate limits per agent, per tool, per hour, enforced in the executor.
  • A kill switch. One flag disables all write tools instantly. New tools default to dry-run mode.

Audit and watch

Log every action: run id, tool, arguments, approver, result. Track approval-queue latency, the share of actions auto-approved, and the rejection rate. A rising rejection rate means the agent's judgement is drifting.

A real-life example

Scenario (illustrative numbers). A subscription-box company's marketing agent can draft and send campaign emails. An engineer debugging a prompt locally runs the agent against production config, and a test email reaches 12,000 customers before anyone notices. Unsubscribes jump by 900 that day.

The fix has four parts. Local and staging runs now load a sandbox mail tool with no production credentials. send_email returns a pending action above 10 recipients. The executor enforces 2,000 emails per hour per agent. And a kill switch in the admin panel disables all write tools. In the next quarter, 340 campaign sends go through approval with a median wait of 6 minutes, and 12 are rejected, 3 of them for wrong audience segments.

Follow-up questions to expect

  • "Won't approvals slow everything down?" — Only for the risky class. Most actions are reads or reversible writes; approvals apply to a small share, and a good preview makes review quick.
  • "Can the model approve its own action?" — No. The approval must come from a person, or a separate policy service, through a channel the model cannot call.
  • "What about prompt injection asking the agent to email everyone?" — The same gates apply; injection can create a proposal, but not an execution.