Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your workflow automation agent can call APIs, databases, browsers, and internal tools. A single bad action could corrupt production systems. How do you sandbox, permission, and monitor tool-using AI agents safely?
What you need to know
Why "untrusted client"
The agent can be wrong (a hallucinated argument), manipulated (a prompt injection in a web page it read), or stuck in a loop. None of these is malicious code, but the effect on production is the same. So design as if the caller cannot be trusted, which is how you would treat a plugin from an unknown vendor.
Layers of control
| Layer | What it does | Example |
|---|---|---|
| Isolation | Tool code runs outside the app process | Containers or microVMs (gVisor, Firecracker); read-only filesystem; no host mounts; egress allowlist |
| Least privilege | Access limited to this task | A 15-minute token that can read one customer's records |
| Policy in code | Checks every call before it runs | Schema validation, value limits, "never production without a flag" |
| Approval by risk | Humans approve irreversible actions | Show the exact rendered call |
| Blast-radius limits | Caps the damage of a mistake | Rate limits per tool, transactions, dry-run mode, kill switch |
| Monitoring | Finds problems fast | Full traces; alerts on unusual tool sequences and repeated failures |
A policy gate
1from pydantic import BaseModel, Field23class RefundArgs(BaseModel):4 order_id: str5 amount_inr: float = Field(gt=0, le=5000) # value limit in the schema67POLICY = {"refund.issue": RefundArgs, "orders.read": None}89def gate(tool: str, args: dict, ctx) -> dict:10 if tool not in POLICY:11 raise PermissionError(f"{tool} is not allowed for this agent")12 schema = POLICY[tool]13 clean = schema(**args).model_dump() if schema else args14 if ctx.env == "production" and tool in ctx.write_tools and not ctx.prod_writes_enabled:15 raise PermissionError("production writes are disabled for this run")16 if ctx.rate_limiter.exceeded(ctx.run_id, tool):17 raise PermissionError("tool rate limit reached")18 return cleanEvery tool call passes through gate before it runs. An unknown tool, an argument outside the schema, a production write without the flag, or too many calls all fail in code, whatever the model says. Pydantic's Field(le=5000) puts the value limit where it cannot be talked around.
- Sandbox tool execution; deny network egress by default.
- Issue scoped, short-lived credentials per task; separate read tools from write tools.
- Gate every call with a policy in code.
- Require approval for irreversible or external actions.
- Limit the blast radius — rate limits, transactions, dry runs, kill switch.
- Trace and alert — every prompt, tool call, argument and result.
A real-life example
Scenario, numbers made up. A platform team's operations agent can run database queries, restart services and edit configuration. During a test, a mis-parsed ticket leads it to run a cleanup query against the production database instead of staging, using the service account it shares with other tools. It deletes 40,000 rows before anyone notices; restoring from backup takes 3 hours.
The rebuild gives the agent per-task credentials that can reach only the environment named in the ticket, a gate that blocks production writes unless a human enables them for that run, and a dry-run step that reports "this would delete 40,000 rows" for approval. Alerts fire on any write above 1,000 rows. In the following quarter the gate blocks 23 risky calls, 4 of them caused by prompt injection in log messages the agent had read.
Follow-up questions to expect
- "Where do MCP servers fit?" — They are just another tool boundary. Give each server its own scoped credentials and run the policy gate on calls before they reach it.
- "Isn't a sandbox too slow?" — MicroVMs start in well under a second, and containers are faster still. Pre-warm a pool for interactive use.
- "How do you detect a compromised agent?" — Unusual sequences, such as reading secrets and then calling an external URL, repeated failed calls, or argument values unlike normal runs.