Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you identify which agent caused a failure?
What you need to know
"Which agent failed?" has two versions: which agent raised an error, and which agent produced a wrong answer. Errors are easy, since they appear in logs with a stack trace. Wrong answers need the contract check from the previous lesson.
Built-in attribution
result = crew.kickoff(inputs=inputs)for i, t in enumerate(result.tasks_output): print(i, t.agent, t.description[:60], len(t.raw))This is often enough in a sequential crew: one task, one agent, one output.
An event listener for the details
CrewAI emits events for crew, task, agent, LLM, tool, memory and guardrail steps. A listener lets you log them with the agent attached:
1from crewai.events import (BaseEventListener, TaskFailedEvent,2 ToolUsageErrorEvent, AgentExecutionCompletedEvent)34class Attribution(BaseEventListener):5 def setup_listeners(self, bus):6 @bus.on(ToolUsageErrorEvent)7 def tool_error(source, event):8 log.warning("tool_error", extra={"run_id": RUN_ID,9 "tool": event.tool_name, "error": str(event.error)})1011 @bus.on(TaskFailedEvent)12 def task_failed(source, event):13 log.error("task_failed", extra={"run_id": RUN_ID,14 "error": str(event.error)})1516 @bus.on(AgentExecutionCompletedEvent)17 def agent_done(source, event):18 log.info("agent_done", extra={"run_id": RUN_ID,19 "agent": event.agent.role})2021attribution = Attribution() # must be created once so it registersA tracing tool gives the same view without code: each agent and tool call appears as a span in a tree.
The hierarchical trap
In a hierarchical crew, the manager writes the instructions each worker receives. If the manager asks the researcher to "check pricing" when the job needed competitor pricing, the researcher's output looks wrong, but the cause is the manager. Always look at the delegation text before blaming the worker. This harder attribution is one reason to prefer sequential crews in production.
A real-life example
A bank's customer-support escalation crew has a classifier, a KYC-lookup agent and a reply writer. A customer receives a reply about a credit card when she wrote about a UPI failure.
The trace shows: the classifier returned category="payments" (correct), the KYC agent called get_customer_products and got a tool error (timeout), then guessed "the customer has a credit card" and carried on. The writer used that guess.
So the culprit is the KYC agent, and the root cause is a tool timeout that was hidden from the logs. Two fixes: the tool now returns a clear "LOOKUP_FAILED" string instead of an empty result, and the task description says "if lookup fails, set products=[] and flag for a human". With a ToolUsageErrorEvent alert, the team found 40 more silent timeouts that week.
Follow-up questions to expect
- "What is the difference between
step_callbackand an event listener?" —step_callbackfires after each agent step and is set on the crew or agent. Event listeners cover many more event types (LLM calls, tools, memory, guardrails, flows) in one place. - "How do you connect logs from one run?" — Generate a
run_idbefore kickoff, pass it ininputsor keep it in the listener, and include it in every log line. - "Can you tell which tool call caused a wrong answer?" — Yes, from the trace: find the tool result that first contains the wrong fact.