Course Content
Agentic AI Patterns
9 sections · 50 lessons
How do you debug agentic workflows using LangChain, AutoGen, or LangGraph?
What you need to know
A framework-neutral method
- Reproduce — same input, pinned model version, same tool responses (record and replay them if you can).
- Find the first bad span — walk forward to the earliest point where reality differs from intent.
- Inspect the prompt as sent — after templating, truncation and message ordering. A variable that rendered empty is a common bug.
- Replay one step — re-run only that model call with the captured context, and vary one thing at a time.
- Add the case to the golden set — so the bug cannot come back silently.
Framework specifics (as of 2026)
- LangGraph: the agent is an explicit graph with typed state. Stream events to watch node transitions. Use a checkpointer to save state after every step; then you can inspect the state at any step, and "time travel" by resuming from an earlier checkpoint with changed state. LangGraph Studio gives a visual view.
- LangChain: since version 1.0, its agents run on the LangGraph runtime, so the same state and checkpoint tools apply. Use callbacks and LangSmith tracing, and look at the rendered prompt and raw tool input and output.
- AutoGen: Microsoft now points new projects to its successor, the Microsoft Agent Framework, and keeps AutoGen in maintenance. In either, log the full inter-agent message history. Most bugs are handoffs: a worker returned prose where JSON was expected, or two agents passed the same task back and forth.
- CrewAI and others: the same idea: turn on verbose logging or tracing and read the task outputs between agents.
Recurring root causes
- Tool schema mismatch, so the model sends a field the tool ignores.
- Tool errors swallowed and returned as empty results.
- Context truncation dropping the task or the rules.
- No step budget, so a loop runs until timeout.
- State not updated, so a node reads a stale value.
A real-life example
An insurance-claims agent built with LangGraph sometimes approved claims without checking prior claims. It happened in about 3% of runs, never in demos.
Debugging:
- The team filtered traces for "approve without
get_prior_claimsspan" and found 41 runs. - The first bad span in each was the router node: it sent the claim down the "simple claim" branch.
- The state at that checkpoint showed
claim_amountas a string with a comma,"48,000". The router'sfloat(claim_amount)failed on the comma, and a broadexceptquietly defaulted to the simple branch. - Replaying from that checkpoint with
claim_amount = 48000sent it to the full-check branch, as the rule says every claim above Rs 25,000 needs a prior-claims check.
The fix was a Pydantic schema on the state, so claim_amount is parsed to a number at entry. The 41 cases joined the golden set. The bug was in plain code, not in the model, which is common.
Follow-up questions to expect
- "How do you debug non-deterministic failures?" — Run the case many times, compare passing and failing traces, and find where they diverge. Record tool responses so only the model varies.
- "What does time travel give you?" — You can resume from any saved step with modified state, so you test a fix at the exact failing point without re-running the whole task.
- "What if you have no tracing?" — Add it first. Debugging an agent without per-step traces is guesswork.