Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you define success criteria for a task?
What you need to know
Level 1: expected_output
This text is what the agent optimises against. Make it checkable:
- Weak: "A good summary."
- Strong: "150-200 words, 3 bullet recommendations, each with a source URL."
Level 2: schema
1from typing import Any2from pydantic import BaseModel, Field3from crewai import Task, TaskOutput45class CompetitorBrief(BaseModel):6 headline: str7 findings: list[str] = Field(min_length=3, max_length=8)8 sources: list[str]910def enough_sources(output: TaskOutput) -> tuple[bool, Any]:11 brief = output.pydantic12 if brief is None or len(brief.sources) < 3:13 return (False, "Cite at least 3 distinct source URLs.")14 return (True, output)1516brief_task = Task(17 description="Write a competitor brief on {competitor}.",18 expected_output="A headline, 3-8 findings and at least 3 source URLs.",19 agent=analyst,20 output_pydantic=CompetitorBrief,21 guardrails=[enough_sources,22 "Every finding must be supported by one of the sources"],23 guardrail_max_retries=2,24)output_pydantic converts the answer into the model; Pydantic enforces types and the min_length rule. The guardrails run in order after that.
Level 3: guardrails
- Function guardrail — takes the
TaskOutput, returns a tuple. Fast, free and deterministic. If you annotate the return type, it must betuple[bool, Any]. - String guardrail — a sentence like "Every finding must be supported by one of the sources". CrewAI turns it into an LLM check using the agent's model. Flexible, but it costs a call and can be wrong.
- On failure, the reason is sent back to the agent and the task runs again. After
guardrail_max_retries(default 3), the task raises an error.
What none of these catch
Tone, usefulness, persuasiveness. For those, keep a golden set of 30 to 50 inputs, run the crew, and score outputs with a rubric (by people or an LLM judge) before releasing changes.
A real-life example
A market-research crew produced competitor briefs for a quick-commerce company's strategy team. The team's complaint: "some briefs have no sources and we can't trust them".
They added the three levels above. In the first week, 14% of first attempts failed enough_sources; after one retry, only 1% still failed and those raised an error that routed the brief to an analyst. The LLM guardrail flagged a further 5% of findings as unsupported, and most were genuinely wrong (for example, a delivery-time figure for the wrong city). Strategy-team complaints about untrusted numbers stopped. Cost per brief rose about 15% because of the LLM guardrail call — which they judged worth it.
Follow-up questions to expect
- "
output_pydanticoroutput_json?" — Both validate against a Pydantic model;output_pydanticgives you the model object,output_jsongives a dict. Useoutput_pydanticwhen your code works with typed fields. - "Function or LLM guardrail?" — Function first, for anything code can check (counts, formats, banned words). LLM guardrails for meaning, such as "claims are supported".
- "What if a guardrail keeps failing?" — That is a signal the task is unclear or impossible with the given inputs. Fix the description or inputs, rather than raising the retry count.