Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 8: Debugging and Observability


What you need to know

The scenario: users report that a crew's reports are sometimes wrong, and the team cannot tell which agent is responsible.

Why crews are hard to debug

A crew run is many LLM calls and tool calls, with outputs feeding each other. The final answer hides all of that. Without records of each step, every bug report becomes "the crew was weird yesterday", and nobody can reproduce it.

What to capture

  1. Tracing — verbose=True locally; in production, OpenTelemetry-based tracing (CrewAI's built-in tracing, or tools such as LangSmith, Langfuse or AgentOps).
  2. Tags — every span carries run_id, crew version, model and prompt version.
  3. Persisted outputs — a task callback writes each task's typed output to your own store.
  4. Per-task metrics — success rate, iterations, tokens, cost and p95 duration per task; error rate per tool.
Python
def save_task_output(output):                       # CrewAI calls this after each task    store.insert("task_outputs", {        "run_id": RUN_ID, "crew_version": CREW_VERSION,        "task": output.name or output.description[:60],        "agent": output.agent,        "raw": output.raw,        "structured": output.pydantic.model_dump() if output.pydantic else None,    })crew = Crew(agents=agents, tasks=tasks, task_callback=save_task_output)

Debug by reduction

Take a failed run. Look at each task's stored output and find the first one that is wrong. Then run that agent alone, with that single task, at temperature 0, on the pinned model, replaying the stored upstream inputs. Most "multi-agent bugs" turn out to be one agent misreading one field.

SymptomLikely place
Final answer has a wrong numberThe first task where that number appears
Runs slow on some inputsOne task's iterations or one tool's latency
Output shape changesMissing schema on a handoff

Regression suite

20–30 real inputs run nightly, asserting schema validity and a few known facts per input. Agent systems drift as models and prompts change, and this suite catches it before users do.

A real-life example

Scenario, numbers made up. A wealth-management crew writes portfolio summaries. About one report in twenty shows the wrong asset allocation, and the team has only final outputs.

After adding run IDs, task callbacks and tracing, they pull five bad runs. In all five, the extractor task's structured output is right, but the analyst's output swaps equity and debt percentages when the source statement lists debt first. Replaying the analyst alone with stored inputs reproduces it every time. A schema with named fields instead of a list fixes it, and a nightly suite of 25 statements, including debt-first layouts, keeps it fixed.

Follow-up questions to expect

  • "Isn't verbose=True enough?" — It helps locally, but console logs are not searchable, joined or kept; production needs structured traces and stored outputs.
  • "What about sensitive data in traces?" — Redact PII before export, restrict access, and set retention like any other customer data.
  • "Which metric do you look at first?" — Per-task failure and iteration counts; failures usually concentrate in one task and one tool.