Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you handle agent failures or incorrect outputs?
What you need to know
Three kinds of failure
- Wrong shape — missing fields, prose instead of JSON.
- Wrong content — a number with no source, a rule broken, a hallucinated fact.
- Runtime failure — a tool times out, an API returns an error, the loop never finishes.
Layer 1: schema
output_pydantic=MyModel makes CrewAI convert the final answer into your model. Downstream code reads task_output.pydantic instead of parsing text.
Layer 2: guardrails
1from typing import Any2from pydantic import BaseModel3from crewai import Task, TaskOutput45class IncomeCheck(BaseModel):6 monthly_income: int7 emi_total: int8 source_pages: list[int]910def has_sources(output: TaskOutput) -> tuple[bool, Any]:11 if output.pydantic is None or not output.pydantic.source_pages:12 return (False, "Give the page number for each figure.")13 return (True, output)1415extract = Task(16 description="Extract monthly income and total EMIs from {file}.",17 expected_output="Income, EMIs and the pages they came from.",18 agent=doc_analyst,19 output_pydantic=IncomeCheck,20 guardrails=[has_sources, "EMI total must not exceed monthly income"],21 guardrail_max_retries=2, # default is 322)A function guardrail takes the TaskOutput and returns a tuple. If you annotate the return type, it must be tuple[bool, Any]. A string guardrail is checked by an LLM. On failure the reason is sent back and the task runs again; after the last retry the task raises an error, which your Flow or calling code must handle.
Layer 3: reviewer
A reviewer agent whose task has context=[extract] and the original document can catch problems code cannot, such as a salary figure read from the wrong month.
Layer 4: runtime limits
max_iter(default 25) — after this, the agent is forced to give its best answer.max_execution_time— a hard time limit in seconds.max_retry_limit(default 2) — how many times the agent retries a task that raised an error.
Layer 5: tools that fail loudly
A tool that returns an empty string on error invites the model to guess. Return a clear message, or return ToolFailure(message=...) from crewai.tools so the framework also records the failure. The agent's tool_failure_policy ("ignore", "warn" or "raise") decides whether that just logs or stops the run.
A real-life example
A loan-document review crew processed 2,000 applications a month. About 4% of outputs had income figures with no page reference, and credit officers could not trust them.
The team added output_pydantic=IncomeCheck and the has_sources guardrail. On first attempts, 4% still failed the check, but after one retry with the reason "Give the page number for each figure", 3.4 of those 4 percentage points passed. The remaining 0.6% — about 12 files a month — raised an error, and the Flow sent them to a human queue instead of guessing. The OCR tool was also changed to return ToolFailure(message="Page 3 unreadable", retryable=False) instead of an empty string, which ended a pattern of "income: Rs 0" outputs.
Follow-up questions to expect
- "Where do you put a retry for a flaky API?" — Inside the tool, with a timeout and a small number of backoff retries. Guardrail retries re-run the whole task, which costs far more.
- "Can a guardrail change the output?" — Yes. Returning
(True, new_value)passes a cleaned value forward, for example trimmed text. - "What happens after the last guardrail retry?" — The task raises an exception; catch it in the Flow or calling code and route the case to a human or a fallback.