CrewAI Multi-Agents

Course Content

CrewAI Multi-Agents

9 sections · 53 lessons

How do you log agent decisions and task outputs?


What you need to know

"Why did the agent do that?" can only be answered if you recorded what it saw and what it chose. CrewAI gives several places to record this, from simple to complete.

The options

  • verbose=True — console output of thoughts and tool calls. For development only; it is noisy and unstructured.
  • output_log_file — True writes logs.txt; a name ending in .json writes JSON.
  • Callbacks — task_callback receives each TaskOutput; step_callback receives each agent step (a tool action or a final answer).
  • before_kickoff / after_kickoff under @CrewBase — log the inputs and the final result of a run.
  • Event listeners — subscribe to events like LLMCallCompletedEvent, ToolUsageFinishedEvent or TaskCompletedEvent.
  • Execution hooks — @before_tool_call and @after_tool_call can log, change or block a tool call.
  • Tracing — tracing=True for CrewAI's built-in traces, or an OpenTelemetry integration (Langfuse, Datadog, MLflow, Braintrust, Opik and others) that records spans with no change to your agents.
Python
import uuidRUN_ID = str(uuid.uuid4())def on_task(output):    log.info("task_done", extra={        "run_id": RUN_ID, "agent": output.agent,        "task": output.description[:80],        "output": redact(output.raw)[:2000],    })crew = Crew(agents=agents, tasks=tasks,            task_callback=on_task,            output_log_file="logs/support_run.json",            verbose=False)

What to log, and what not to

Log per event: run_id, crew name, task, agent role, tool name and arguments, latency, token counts, and a shortened output. Do not send full prompts with customer data or raw tool responses with personal details (phone numbers, account numbers, Aadhaar or PAN) to an external vendor. Redact in the callback, because once it is in a vendor's log store it is hard to remove.

Logs you will actually use

  • Steps per run — a jump means an agent is looping.
  • Tool errors per tool — finds silent failures.
  • Guardrail failures per task — shows which task contract is weak.
  • Tokens per run — early warning for cost problems.

A real-life example

A fintech's customer-support escalation crew handles 5,000 tickets a day. Its first version only had verbose=True, so when a customer complained that the bot "refused to unblock my card", nobody could reconstruct the run.

The team added a task_callback, an event listener for tool calls and guardrail results, and Langfuse tracing. Phone numbers and card numbers are masked by a redact() function before logging. The next complaint took five minutes to explain: the get_card_status tool returned BLOCKED_BY_BANK (a fraud hold), and the writer agent correctly refused. The logs also showed that 7% of runs used more than 15 agent steps; those were all one ticket type with a vague task, which the team rewrote.

Follow-up questions to expect

  • "Do you log the model's chain of thought?" — I log the agent's steps and tool calls. For reasoning models, the hidden reasoning may not be available, and the tool calls are the more reliable record anyway.
  • "How long do you keep logs?" — Full payloads for a short period (say 30 days) for debugging; metrics and redacted summaries longer.
  • "Does logging slow the crew down?" — Callbacks run inline, so keep them light and send heavy work to a queue.