Course Content
CrewAI Multi-Agents
9 sections · 53 lessons
How do you debug incorrect outputs in a CrewAI workflow?
What you need to know
A crew is a chain. When the final report is wrong, the mistake usually happened earlier and was passed along. The writer faithfully wrote up a wrong number the analyst gave it. So the job is to find the first broken link.
What you can look at
result.tasks_output— a list ofTaskOutputobjects, one per task, in order. Each hasdescription,agent(the role),raw, andpydanticif you set a schema.verbose=True— prints each agent's thoughts, tool calls and tool results to the console. Good while developing.output_log_file— saves the run log to a.txtor.jsonfile, for runs you cannot watch.- Tracing —
tracing=Truesends traces to CrewAI's own tracing view, or you can use an OpenTelemetry integration such as Langfuse, Datadog, MLflow or Opik. A trace shows every LLM call and tool call with timings.
A debugging procedure
- Capture — get the full trace of the bad run, with inputs.
- Locate — walk
tasks_outputand find the earliest output that fails its ownexpected_output. - Classify — is it a bad tool result, missing
context, an unclear task description, or a real model error? - Isolate — rerun just that task with the exact upstream input, in a one-agent crew.
- Fix and resume — change the narrowest thing, then resume from that task instead of rerunning everything.
In practice, the first three causes are much more common than the model simply being wrong. A useful test: could a careful human have produced the expected output from that task description and those inputs? If not, the task is the bug.
Resuming instead of rerunning
crewai log-tasks-outputs # list task IDs from the latest kickoffcrewai replay -t <task_id> # rerun from that task, reusing earlier outputscrew.replay(task_id=...) does the same in Python. Newer CrewAI versions also have checkpointing: Crew(..., checkpoint=True) saves state to .checkpoints/ after every task, and Crew.from_checkpoint(path).kickoff() resumes from it. Replay is for debugging the last run; checkpoints are for recovering production runs.
A real-life example
A market-research crew reports that the Indian electric two-wheeler market is "₹8 lakh crore". The number is far too high, and the client notices.
Walking tasks_output:
| Task | Output | Verdict |
|---|---|---|
| 1. Search | 12 sources, including one on all two-wheelers | looks fine |
| 2. Extract figures | "market size: ₹8 lakh crore" from that source | first wrong |
| 3. Analyse | growth rates built on the wrong base | downstream |
| 4. Write | a well-written wrong report | downstream |
Cause: the extract task said "find the market size" without saying electric only. The fix is one line in task 2's description plus a segment field in its schema. The engineer replays from task 2 (four calls, about ₹6) instead of rerunning the full crew with web searches (about ₹40), and tries three wording changes in ten minutes.
Follow-up questions to expect
- "Would you switch to a bigger model first?" — No. I first check whether the task and inputs were clear. A bigger model hides a vague task for a while, at a higher price.
- "How do you debug a hierarchical crew?" — Same idea, but also log the manager's delegations, because the manager wrote the sub-task text the worker received.
- "What if the output is only wrong sometimes?" — Run the same input several times and compare traces. The step where the runs split is where the prompt or tool is ambiguous.