CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you test an entire crew end-to-end?


The built-in check against your own gatecrewai test• Runs the crew n times, default 2• Generic judge scores tasks 1 to 10• Evaluator supports OpenAI models only• Good as a quick smoke testYour golden set• 20 to 50 real inputs, hard ones included• Checks facts, schema and tool calls• Compares to the last good version• Blocks a release on a real drop
The built-in judge does not know your business rules, so it can pass a report that has lost its sources.

What you need to know

A crew's output changes from run to run, and a single run proves little. So an end-to-end test asks a different question: is this version at least as good as the last one, on the cases we care about?

What to check on each golden case

  • Completion — no exception, finished within, say, 3 minutes and 60,000 tokens.
  • Schema — result.pydantic validates.
  • Behaviour — required tools were called (from your logs or hooks); for example, the policy lookup always runs.
  • Quality — rubric scores (accuracy, completeness, tone) from a person or an LLM judge, compared to the baseline.
  • Cost — result.token_usage within about 20% of the baseline. Token growth is often the first sign of a new loop.

The built-in crewai test

Bash
crewai test -n 3 -m gpt-4o-mini

This runs the crew n times (default 2) and uses an evaluator model (default gpt-4o-mini) to score each task and the crew from 1 to 10, with execution times. The docs say it currently supports only OpenAI models for the evaluator. It is quick, but the judge does not know your business rules, so treat it as a smoke test.

Do not confuse it with crewai train, which runs the crew with a human giving feedback after each iteration and saves that feedback to a .pkl file the agents use later. Training changes behaviour; testing measures it.

Keeping the results meaningful

  • Pin model versions. A silent provider update can move scores and look like your regression.
  • Compare to a baseline. Store scores per case and alert on drops, not on an absolute bar.
  • Run on a schedule. Each full run costs real money and takes minutes, so nightly and pre-release is usually right.

A real-life example

A media company's daily news-digest flow runs at 6 am. Its golden set is 30 past days of news, each with a checklist: "must mention the RBI rate decision", "no section over 130 words", "every figure has a source".

The nightly test runs all 30 days, costing about ₹900. One evening a developer shortened the fact-gathering task's description. Schema and completion checks passed, but "every figure has a source" dropped from 29 of 30 days to 21 of 30. The release was blocked, and the diff showed the removed line: "include the URL for each figure".

Follow-up questions to expect

  • "Can you use crewai test in CI?" — You can, but it costs money on each run and its scores are noisy. I use it as a quick check and gate releases on my own golden set.
  • "How big should the golden set be?" — Start with 20 cases that cover the main input types and every past incident. Add a case every time something breaks in production.
  • "How do you test a Flow with branches?" — Include cases for each branch, and assert which branch ran from the flow state.