CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you test individual agents in isolation?


What you need to know

An agent's output is not the same every run, so "assert output == expected" rarely works for text. But much of an agent can be tested exactly: its tools, its output shape, and whether it used the right tool.

Layer 1: tools, with no LLM

Tools are ordinary Python classes. Call _run() directly with test data. Most real bugs (wrong API field, bad date parsing) live here.

Python
def test_policy_lookup_finds_clause():    tool = PolicyLookupTool(db=fake_db)    assert "4.2(b)" in tool._run(policy_id="HOME-2231", topic="water")

Layer 2: structure, with a real or stubbed model

Python
from pydantic import BaseModelclass ClaimFacts(BaseModel):    claim_id: str    amount_inr: int    incident_date: strdef test_intake_agent_extracts_fields():    out = intake_agent.kickoff(        f"Extract the claim facts:\n{FIXTURE_CLAIM_TEXT}",        response_format=ClaimFacts,    )    facts = out.pydantic    assert facts.claim_id == "CLM-1001"    assert facts.amount_inr == 380000

Agent.kickoff() runs one agent without a crew and returns a LiteAgentOutput with raw, pydantic and usage_metrics. Under @CrewBase, MyCrew().intake_agent() returns the configured agent, so the test uses the real YAML config. Use a low temperature, and a small model if you only check structure.

To check that a tool was called, record calls with a @before_tool_call hook or a fake tool that counts its calls.

Layer 3: quality, on a golden set

For judgement, like "is the summary accurate and polite", keep 20–50 real examples with notes on what a good answer contains, and score outputs with a rubric (a person or an LLM judge). Compare against the last good version, because the scores are noisy.

LayerSpeedCostWhen
Toolsmillisecondsfreeevery commit
Structuresecondssmallevery commit or PR
Qualityminutesreal moneynightly or before release

A real-life example

An insurance-claims review crew has an intake agent that reads claim forms. A change to its backstory made it write amounts as "₹3.8 lakh" instead of 380000, which broke the policy agent downstream.

The team added a structure test with 12 fixture claims, including a handwritten scan and a form in Hindi. It runs in CI in about 40 seconds with a small model and costs a few rupees. The test failed on the next backstory change before it was merged. The tool tests (policy lookup, date parser) run in under a second and caught a timezone bug that had been blamed on "the model".

Follow-up questions to expect

  • "Can you mock the LLM?" — Yes, for pure wiring tests, but then you are not testing the agent's behaviour. I use a real, cheap model for structure tests and mocks only for tool and flow logic.
  • "How do you make tests less flaky?" — Assert on schema fields and tool calls, not wording; use low temperature; and allow a pass rate (for example 9 of 10) for quality tests.
  • "What about crewai test?" — That tests a whole crew, not one agent. See the next lesson.