Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you test individual agents in isolation?
What you need to know
An agent's output is not the same every run, so "assert output == expected" rarely works for text. But much of an agent can be tested exactly: its tools, its output shape, and whether it used the right tool.
Layer 1: tools, with no LLM
Tools are ordinary Python classes. Call _run() directly with test data. Most real bugs (wrong API field, bad date parsing) live here.
def test_policy_lookup_finds_clause(): tool = PolicyLookupTool(db=fake_db) assert "4.2(b)" in tool._run(policy_id="HOME-2231", topic="water")Layer 2: structure, with a real or stubbed model
1from pydantic import BaseModel23class ClaimFacts(BaseModel):4 claim_id: str5 amount_inr: int6 incident_date: str78def test_intake_agent_extracts_fields():9 out = intake_agent.kickoff(10 f"Extract the claim facts:\n{FIXTURE_CLAIM_TEXT}",11 response_format=ClaimFacts,12 )13 facts = out.pydantic14 assert facts.claim_id == "CLM-1001"15 assert facts.amount_inr == 380000Agent.kickoff() runs one agent without a crew and returns a LiteAgentOutput with raw, pydantic and usage_metrics. Under @CrewBase, MyCrew().intake_agent() returns the configured agent, so the test uses the real YAML config. Use a low temperature, and a small model if you only check structure.
To check that a tool was called, record calls with a @before_tool_call hook or a fake tool that counts its calls.
Layer 3: quality, on a golden set
For judgement, like "is the summary accurate and polite", keep 20–50 real examples with notes on what a good answer contains, and score outputs with a rubric (a person or an LLM judge). Compare against the last good version, because the scores are noisy.
| Layer | Speed | Cost | When |
|---|---|---|---|
| Tools | milliseconds | free | every commit |
| Structure | seconds | small | every commit or PR |
| Quality | minutes | real money | nightly or before release |
A real-life example
An insurance-claims review crew has an intake agent that reads claim forms. A change to its backstory made it write amounts as "₹3.8 lakh" instead of 380000, which broke the policy agent downstream.
The team added a structure test with 12 fixture claims, including a handwritten scan and a form in Hindi. It runs in CI in about 40 seconds with a small model and costs a few rupees. The test failed on the next backstory change before it was merged. The tool tests (policy lookup, date parser) run in under a second and caught a timezone bug that had been blamed on "the model".
Follow-up questions to expect
- "Can you mock the LLM?" — Yes, for pure wiring tests, but then you are not testing the agent's behaviour. I use a real, cheap model for structure tests and mocks only for tool and flow logic.
- "How do you make tests less flaky?" — Assert on schema fields and tool calls, not wording; use low temperature; and allow a pass rate (for example 9 of 10) for quality tests.
- "What about
crewai test?" — That tests a whole crew, not one agent. See the next lesson.