Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you prevent agents from hallucinating when using tools?
What you need to know
Where tool hallucinations come from
- Silent failures. The tool returned
""orNone, and the model filled the gap. - Truncated results. A long result was cut off, and the model guessed the rest.
- Tool not called. A weak tool description meant the model answered from its own knowledge.
- Loop exhaustion. After many failed calls, the agent hit
max_iterand gave its "best answer" from memory.
Five defences
- Fail loudly — return "ERROR: no results for X" or a
ToolFailure; say "do not guess" in the message. - Return citable data — every result carries a URL, document ID or page number.
- Require citations — in
expected_outputand inoutput_pydantic(for examplesources: list[str]). - Check citations with a guardrail — every claim has a source, and every source was actually returned by a tool.
- Verify — for high stakes, a separate reviewer task gets the draft and the raw tool outputs and flags unsupported claims.
Python
1from typing import Any2from pydantic import BaseModel3from crewai import TaskOutput45class Claim(BaseModel):6 text: str7 source_url: str89class Findings(BaseModel):10 claims: list[Claim]1112def sources_were_retrieved(output: TaskOutput) -> tuple[bool, Any]:13 if output.pydantic is None:14 return (False, "Return findings in the required structure.")15 retrieved = load_urls_seen_this_run() # e.g. logged by the search tool16 missing = [c.text for c in output.pydantic.claims17 if c.source_url not in retrieved]18 if missing:19 return (False, f"These claims cite URLs never retrieved: {missing[:3]}")20 return (True, output)load_urls_seen_this_run stands for your own record of URLs the search tool returned. The guardrail catches a subtle failure: a real-looking URL the model made up.
A real-life example
A market-research crew for a fintech's strategy team produced a report saying a competitor's personal-loan rate was "10.49% p.a." The number came from nowhere: the search tool had timed out and returned an empty string, and the agent filled the gap with a plausible figure.
The team made three changes:
- The search tool now returns
ToolFailure(message="Search timed out; no data. Do not estimate."). - Findings became a
Findingsmodel with asource_urlper claim. - The
sources_were_retrievedguardrail checks every URL against the search tool's log for the run.
Over the next 100 reports, the guardrail blocked 7 drafts with made-up URLs; 6 passed after one retry with real sources, and 1 was sent to an analyst. Reports now say "rate not found" when search fails, which the strategy team prefers to a confident wrong number.
Follow-up questions to expect
- "Does a lower temperature stop hallucination?" — It reduces randomness, not fabrication. A model at temperature 0 can still invent a value when a tool returns nothing.
- "Can an LLM judge catch hallucinations?" — Partly: a reviewer comparing claims to raw sources catches many, but it can be wrong too, so combine it with code checks like the URL guardrail.
- "What about hallucinated tool arguments?" — Schema validation, regex patterns and "not found" responses handle invented IDs.