Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you integrate external APIs or databases as tools?
What you need to know
A database tool done properly
1from pydantic import BaseModel, Field2from crewai.tools import BaseTool, ToolFailure34class LoanLookupInput(BaseModel):5 application_id: str = Field(..., pattern=r"^APP-\d{6}$",6 description="Application ID, e.g. APP-104233")78class LoanLookupTool(BaseTool):9 name: str = "loan_application_lookup"10 description: str = ("Fetch applicant name, loan amount, tenure and document "11 "list for one application ID. Read-only.")12 args_schema: type[BaseModel] = LoanLookupInput1314 def _run(self, application_id: str):15 try:16 row = db.fetch_one(LOAN_QUERY, {"id": application_id}, timeout=5)17 except TimeoutError:18 return ToolFailure(message="Loan database timed out. Do not guess values.",19 retryable=True)20 if row is None:21 return f"No application found with ID {application_id}."22 return row.model_dump_json() # a few fields, not the whole tabledb, LOAN_QUERY and row stand for your own read-only data layer. The explanation:
- Schema — the
patternrejects anything that is not a valid ID before your code runs. - Fixed query — the model supplies a value, never SQL.
- Timeout — five seconds, so a hung database does not hang the agent.
- Honest errors —
ToolFailuresends text the model can read and marks the call as failed. The agent'stool_failure_policythen decides:"warn"(the default) records it and continues,"raise"stops the run. - "Not found" is not an error — it is a valid, useful answer.
Rules for any external API
| Rule | Why |
|---|---|
| Least-privilege credentials | a tricked agent cannot do more than the key allows |
| Parameterised calls only | stops SQL injection and URL tampering |
| Timeouts and limited retries in the tool | cheaper than re-running the whole task |
| Compact outputs (fields, top N rows) | every result is re-sent on each later step |
| Idempotency keys for writes | agent retries must not create duplicates |
| Log each call with the run ID | you can trace which agent did what |
Writes need extra care
Reading is safe to retry; writing is not. If a tool creates a CRM note or sends an email, give it an idempotency key (for example, run ID plus lead ID), or keep the write outside the agent: the crew decides, and your Flow code performs the action.
A real-life example
A loan-document review crew at an NBFC first used a generic "SQL query" tool so the agent could "look up anything". In testing, the agent ran a SELECT * on the applications table and pulled 20,000 rows into its context, which failed with a context-length error; another run built a query with a string-concatenated customer name.
The team replaced it with three narrow tools — loan_application_lookup, bank_statement_summary, bureau_score_lookup — each with a regex-checked ID, a read-only database user and a 5-second timeout. Context errors disappeared, the average tool result shrank from about 6,000 tokens to about 250, and the security review passed because no model-written SQL reached the database.
Follow-up questions to expect
- "Why not let the model write SQL for flexibility?" — Because it can read data it should not, write when it should not, and produce slow queries. If you need ad-hoc analytics, use a read-only replica, a restricted view and query limits — and still review the risk.
- "Where do retries go — tool or task?" — Network retries in the tool, with backoff. Task retries re-run the whole LLM loop and cost far more.
- "How do you handle secrets?" — Environment variables or a secret manager read by the tool; never in prompts, backstories or task text.