Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your workflow automation agent can call APIs, databases, browsers, and internal tools. A single bad action could corrupt production systems. How do you sandbox, permission, and monitor tool-using AI agents safely?


What you need to know

Why "untrusted client"

The agent can be wrong (a hallucinated argument), manipulated (a prompt injection in a web page it read), or stuck in a loop. None of these is malicious code, but the effect on production is the same. So design as if the caller cannot be trusted, which is how you would treat a plugin from an unknown vendor.

Layers of control

LayerWhat it doesExample
IsolationTool code runs outside the app processContainers or microVMs (gVisor, Firecracker); read-only filesystem; no host mounts; egress allowlist
Least privilegeAccess limited to this taskA 15-minute token that can read one customer's records
Policy in codeChecks every call before it runsSchema validation, value limits, "never production without a flag"
Approval by riskHumans approve irreversible actionsShow the exact rendered call
Blast-radius limitsCaps the damage of a mistakeRate limits per tool, transactions, dry-run mode, kill switch
MonitoringFinds problems fastFull traces; alerts on unusual tool sequences and repeated failures

A policy gate

Python
from pydantic import BaseModel, Fieldclass RefundArgs(BaseModel):    order_id: str    amount_inr: float = Field(gt=0, le=5000)     # value limit in the schemaPOLICY = {"refund.issue": RefundArgs, "orders.read": None}def gate(tool: str, args: dict, ctx) -> dict:    if tool not in POLICY:        raise PermissionError(f"{tool} is not allowed for this agent")    schema = POLICY[tool]    clean = schema(**args).model_dump() if schema else args    if ctx.env == "production" and tool in ctx.write_tools and not ctx.prod_writes_enabled:        raise PermissionError("production writes are disabled for this run")    if ctx.rate_limiter.exceeded(ctx.run_id, tool):        raise PermissionError("tool rate limit reached")    return clean

Every tool call passes through gate before it runs. An unknown tool, an argument outside the schema, a production write without the flag, or too many calls all fail in code, whatever the model says. Pydantic's Field(le=5000) puts the value limit where it cannot be talked around.

  1. Sandbox tool execution; deny network egress by default.
  2. Issue scoped, short-lived credentials per task; separate read tools from write tools.
  3. Gate every call with a policy in code.
  4. Require approval for irreversible or external actions.
  5. Limit the blast radius — rate limits, transactions, dry runs, kill switch.
  6. Trace and alert — every prompt, tool call, argument and result.

A real-life example

Scenario, numbers made up. A platform team's operations agent can run database queries, restart services and edit configuration. During a test, a mis-parsed ticket leads it to run a cleanup query against the production database instead of staging, using the service account it shares with other tools. It deletes 40,000 rows before anyone notices; restoring from backup takes 3 hours.

The rebuild gives the agent per-task credentials that can reach only the environment named in the ticket, a gate that blocks production writes unless a human enables them for that run, and a dry-run step that reports "this would delete 40,000 rows" for approval. Alerts fire on any write above 1,000 rows. In the following quarter the gate blocks 23 risky calls, 4 of them caused by prompt injection in log messages the agent had read.

Follow-up questions to expect

  • "Where do MCP servers fit?" — They are just another tool boundary. Give each server its own scoped credentials and run the policy gate on calls before they reach it.
  • "Isn't a sandbox too slow?" — MicroVMs start in well under a second, and containers are faster still. Pre-warm a pool for interactive use.
  • "How do you detect a compromised agent?" — Unusual sequences, such as reading secrets and then calling an external URL, repeated failed calls, or argument values unlike normal runs.