CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you debug incorrect outputs in a CrewAI workflow?


Walking tasks_output for the wrong market sizesearchextractanalysewrite0123first wrong:all two-wheelerswell written,still wrongReplaying from extract cost about Rs 6 a try instead of Rs 40 for the full crew.
Everything after the first broken contract is downstream damage, so fixing the last agent's prompt changes nothing.

What you need to know

A crew is a chain. When the final report is wrong, the mistake usually happened earlier and was passed along. The writer faithfully wrote up a wrong number the analyst gave it. So the job is to find the first broken link.

What you can look at

  • result.tasks_output — a list of TaskOutput objects, one per task, in order. Each has description, agent (the role), raw, and pydantic if you set a schema.
  • verbose=True — prints each agent's thoughts, tool calls and tool results to the console. Good while developing.
  • output_log_file — saves the run log to a .txt or .json file, for runs you cannot watch.
  • Tracing — tracing=True sends traces to CrewAI's own tracing view, or you can use an OpenTelemetry integration such as Langfuse, Datadog, MLflow or Opik. A trace shows every LLM call and tool call with timings.

A debugging procedure

  1. Capture — get the full trace of the bad run, with inputs.
  2. Locate — walk tasks_output and find the earliest output that fails its own expected_output.
  3. Classify — is it a bad tool result, missing context, an unclear task description, or a real model error?
  4. Isolate — rerun just that task with the exact upstream input, in a one-agent crew.
  5. Fix and resume — change the narrowest thing, then resume from that task instead of rerunning everything.

In practice, the first three causes are much more common than the model simply being wrong. A useful test: could a careful human have produced the expected output from that task description and those inputs? If not, the task is the bug.

Resuming instead of rerunning

Bash
crewai log-tasks-outputs     # list task IDs from the latest kickoffcrewai replay -t <task_id>   # rerun from that task, reusing earlier outputs

crew.replay(task_id=...) does the same in Python. Newer CrewAI versions also have checkpointing: Crew(..., checkpoint=True) saves state to .checkpoints/ after every task, and Crew.from_checkpoint(path).kickoff() resumes from it. Replay is for debugging the last run; checkpoints are for recovering production runs.

A real-life example

A market-research crew reports that the Indian electric two-wheeler market is "₹8 lakh crore". The number is far too high, and the client notices.

Walking tasks_output:

TaskOutputVerdict
1. Search12 sources, including one on all two-wheelerslooks fine
2. Extract figures"market size: ₹8 lakh crore" from that sourcefirst wrong
3. Analysegrowth rates built on the wrong basedownstream
4. Writea well-written wrong reportdownstream

Cause: the extract task said "find the market size" without saying electric only. The fix is one line in task 2's description plus a segment field in its schema. The engineer replays from task 2 (four calls, about ₹6) instead of rerunning the full crew with web searches (about ₹40), and tries three wording changes in ten minutes.

Follow-up questions to expect

  • "Would you switch to a bigger model first?" — No. I first check whether the task and inputs were clear. A bigger model hides a vague task for a while, at a higher price.
  • "How do you debug a hierarchical crew?" — Same idea, but also log the manager's delegations, because the manager wrote the sub-task text the worker received.
  • "What if the output is only wrong sometimes?" — Run the same input several times and compare traces. The step where the runs split is where the prompt or tool is ambiguous.