CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you prevent agents from hallucinating when using tools?


How an invented interest rate got into a reportSearch tooltimes outReturns anempty stringAgentfills the gapReport states10.49% p.a.Fix: ToolFailure on timeout, a source URL per claim, and a guardrail that checks each URL was retrieved.
Most tool hallucinations start as a silent tool failure, so the cure is an honest error, not a stricter prompt.

What you need to know

Where tool hallucinations come from

  • Silent failures. The tool returned "" or None, and the model filled the gap.
  • Truncated results. A long result was cut off, and the model guessed the rest.
  • Tool not called. A weak tool description meant the model answered from its own knowledge.
  • Loop exhaustion. After many failed calls, the agent hit max_iter and gave its "best answer" from memory.

Five defences

  1. Fail loudly — return "ERROR: no results for X" or a ToolFailure; say "do not guess" in the message.
  2. Return citable data — every result carries a URL, document ID or page number.
  3. Require citations — in expected_output and in output_pydantic (for example sources: list[str]).
  4. Check citations with a guardrail — every claim has a source, and every source was actually returned by a tool.
  5. Verify — for high stakes, a separate reviewer task gets the draft and the raw tool outputs and flags unsupported claims.
Python
from typing import Anyfrom pydantic import BaseModelfrom crewai import TaskOutputclass Claim(BaseModel):    text: str    source_url: strclass Findings(BaseModel):    claims: list[Claim]def sources_were_retrieved(output: TaskOutput) -> tuple[bool, Any]:    if output.pydantic is None:        return (False, "Return findings in the required structure.")    retrieved = load_urls_seen_this_run()        # e.g. logged by the search tool    missing = [c.text for c in output.pydantic.claims               if c.source_url not in retrieved]    if missing:        return (False, f"These claims cite URLs never retrieved: {missing[:3]}")    return (True, output)

load_urls_seen_this_run stands for your own record of URLs the search tool returned. The guardrail catches a subtle failure: a real-looking URL the model made up.

A real-life example

A market-research crew for a fintech's strategy team produced a report saying a competitor's personal-loan rate was "10.49% p.a." The number came from nowhere: the search tool had timed out and returned an empty string, and the agent filled the gap with a plausible figure.

The team made three changes:

  • The search tool now returns ToolFailure(message="Search timed out; no data. Do not estimate.").
  • Findings became a Findings model with a source_url per claim.
  • The sources_were_retrieved guardrail checks every URL against the search tool's log for the run.

Over the next 100 reports, the guardrail blocked 7 drafts with made-up URLs; 6 passed after one retry with real sources, and 1 was sent to an analyst. Reports now say "rate not found" when search fails, which the strategy team prefers to a confident wrong number.

Follow-up questions to expect

  • "Does a lower temperature stop hallucination?" — It reduces randomness, not fabrication. A model at temperature 0 can still invent a value when a tool returns nothing.
  • "Can an LLM judge catch hallucinations?" — Partly: a reviewer comparing claims to raw sources catches many, but it can be wrong too, so combine it with code checks like the URL guardrail.
  • "What about hallucinated tool arguments?" — Schema validation, regex patterns and "not found" responses handle invented IDs.