Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 7: Tool Integration Safety
What you need to know
The scenario: a crew is getting tools that write to the CRM, send emails and run code. The security team wants to know what stops it from doing damage.
Why tools are the risk
The model decides which tool to call and with what arguments. A confused model, or one manipulated by text in a web page it read, can pass harmful arguments. So the tool must be safe whatever arguments arrive.
The controls
| Control | What it prevents |
|---|---|
| Read-only credentials by default; narrow write tools | A single tool with unlimited reach |
Argument validation in code (args_schema, enums, ranges) | Harmful or malformed inputs |
| Idempotency keys on writes | Duplicate actions on retry |
| Approval for destructive or external actions, from a policy table | The model deciding it doesn't need permission |
| Sandboxing for code tools: container, no network, CPU and time limits | Arbitrary code reaching your systems |
| Per-run budgets: tool calls, spend, emails | Runaway crews |
| Tool output treated as data, clearly delimited | Indirect prompt injection from fetched content |
| Logging every call with arguments and result | Invisible misuse |
1from typing import Literal2from pydantic import BaseModel, EmailStr, Field3from crewai.tools import BaseTool45class SendEmailArgs(BaseModel):6 to: EmailStr7 template: Literal["renewal_reminder", "meeting_followup"] # no free-form bodies8 customer_id: str = Field(pattern=r"^CUST-\d{6}$")910class SendEmail(BaseTool):11 name: str = "send_email"12 description: str = "Send an approved template email to one existing customer."13 args_schema: type[BaseModel] = SendEmailArgs1415 def _run(self, to: str, template: str, customer_id: str) -> str:16 if crm.email_of(customer_id) != to:17 return "Refused: address does not match this customer."18 if run_budget.emails_sent >= 20:19 return "Refused: email budget for this run is used up."20 key = f"{RUN_ID}:{customer_id}:{template}"21 if DRY_RUN:22 dry_run_log.append({"to": to, "template": template, "key": key})23 return "Dry run: email logged, not sent."24 return mailer.send(to, template, idempotency_key=key)Dry-run mode
Every write tool can log its payload instead of acting. Run the full crew in dry-run in CI and compare the payloads with expectations. That is what catches "it would have emailed 12,000 people" before it happens.
A real-life example
Scenario, numbers made up. A sales-ops crew is given a generic http_request tool "for flexibility". In testing, a scraped competitor page contains hidden text telling the agent to post data to an external URL, and the agent tries.
The team removes the generic tool and adds three narrow tools: read CRM account, create a follow-up task, and send one of two email templates to an existing customer. Emails need approval above five per run. A dry-run CI job runs 30 scenarios nightly. Over the next quarter, the logs show 140 refused calls (mostly budget and address mismatches) and zero unintended sends.
Follow-up questions to expect
- "Isn't argument validation the model's job?" — No. Prompts guide the model; only code can enforce limits.
- "How do you handle a tool that must fetch arbitrary web pages?" — A domain allowlist, size and time limits, stripping scripts, and treating the content as data that never triggers write tools directly.
- "Who decides which actions need approval?" — A policy table in code, reviewed like any access-control change.