CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you define success criteria for a task?


What you need to know

Level 1: expected_output

This text is what the agent optimises against. Make it checkable:

  • Weak: "A good summary."
  • Strong: "150-200 words, 3 bullet recommendations, each with a source URL."

Level 2: schema

Python
from typing import Anyfrom pydantic import BaseModel, Fieldfrom crewai import Task, TaskOutputclass CompetitorBrief(BaseModel):    headline: str    findings: list[str] = Field(min_length=3, max_length=8)    sources: list[str]def enough_sources(output: TaskOutput) -> tuple[bool, Any]:    brief = output.pydantic    if brief is None or len(brief.sources) < 3:        return (False, "Cite at least 3 distinct source URLs.")    return (True, output)brief_task = Task(    description="Write a competitor brief on {competitor}.",    expected_output="A headline, 3-8 findings and at least 3 source URLs.",    agent=analyst,    output_pydantic=CompetitorBrief,    guardrails=[enough_sources,                "Every finding must be supported by one of the sources"],    guardrail_max_retries=2,)

output_pydantic converts the answer into the model; Pydantic enforces types and the min_length rule. The guardrails run in order after that.

Level 3: guardrails

  • Function guardrail — takes the TaskOutput, returns a tuple. Fast, free and deterministic. If you annotate the return type, it must be tuple[bool, Any].
  • String guardrail — a sentence like "Every finding must be supported by one of the sources". CrewAI turns it into an LLM check using the agent's model. Flexible, but it costs a call and can be wrong.
  • On failure, the reason is sent back to the agent and the task runs again. After guardrail_max_retries (default 3), the task raises an error.

What none of these catch

Tone, usefulness, persuasiveness. For those, keep a golden set of 30 to 50 inputs, run the crew, and score outputs with a rubric (by people or an LLM judge) before releasing changes.

A real-life example

A market-research crew produced competitor briefs for a quick-commerce company's strategy team. The team's complaint: "some briefs have no sources and we can't trust them".

They added the three levels above. In the first week, 14% of first attempts failed enough_sources; after one retry, only 1% still failed and those raised an error that routed the brief to an analyst. The LLM guardrail flagged a further 5% of findings as unsupported, and most were genuinely wrong (for example, a delivery-time figure for the wrong city). Strategy-team complaints about untrusted numbers stopped. Cost per brief rose about 15% because of the LLM guardrail call — which they judged worth it.

Follow-up questions to expect

  • "output_pydantic or output_json?" — Both validate against a Pydantic model; output_pydantic gives you the model object, output_json gives a dict. Use output_pydantic when your code works with typed fields.
  • "Function or LLM guardrail?" — Function first, for anything code can check (counts, formats, banned words). LLM guardrails for meaning, such as "claims are supported".
  • "What if a guardrail keeps failing?" — That is a signal the task is unclear or impossible with the given inputs. Fix the description or inputs, rather than raising the retry count.