Course Content
LangChain Mastery
7 sections · 109 lessons
How do you debug LangChain agent decision-making?
What you need to know
Classify the failure first
When an agent gives a wrong answer, find which step went wrong. There are only a few possibilities:
| Symptom in the trace | Likely cause |
|---|---|
| Called the wrong tool | Overlapping or vague descriptions |
| Right tool, wrong arguments | Missing field descriptions or examples in the schema |
| Right result, wrong answer | Tool output too long or unclear; model misread it |
| No tool call, guessed the answer | System prompt does not require tools for facts |
| Same call repeated many times | Error message gives no guidance; no call limit |
Seeing the steps locally
1result = agent.invoke({"messages": [{"role": "user", "content": question}]})2for m in result["messages"]:3 if m.type == "ai" and m.tool_calls:4 for c in m.tool_calls:5 print("CALL ", c["name"], c["args"])6 elif m.type == "tool":7 print("RESULT", m.name, str(m.content)[:200])8 elif m.type == "ai":9 print("ANSWER", m.content[:200])The message list is the whole story of the run. For live output while it runs, use for chunk in agent.stream(inputs, stream_mode="updates"): print(chunk), which prints each node's update (the model's decision, then the tool results) as it happens. set_debug(True) from langchain_core.globals prints everything, but the volume makes it hard to read.
Tracing with LangSmith
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=support-bot-devWith these set, every run is recorded with no code change. The trace shows what a print statement cannot: the full prompt including the tool schemas, token counts and latency per step, and which middleware changed the request. You can open a bad run, edit the prompt in the playground, and re-run that exact step.
Fixing, in order of leverage
- Tool descriptions — add "use for X, do not use for Y", and an example argument. This fixes most wrong choices.
- Tool count — merge or remove overlapping tools, or filter tools per request with middleware.
- Tool output — return a short, labelled result instead of raw JSON.
- System prompt — state the policy: "Always check the order tool before answering about delivery."
- Model — only then try a stronger model.
Make it a test
Build 30 to 100 real questions with the expected tool sequence, such as ["get_order", "track_shipment"]. Run them on every prompt, tool or model change and assert on the tools called, not only on the final text. LangSmith datasets and evaluators can run this; a plain pytest loop over result["messages"] also works.
A real-life example
A law firm's contract-search assistant keeps answering "Which contracts with Acme auto-renew?" with "I found no such contracts" — but they exist.
The trace shows the agent calling semantic_search("Acme auto-renew"), getting 5 chunks from unrelated contracts, and giving up. It never called find_contracts(counterparty="Acme"), which would have found all 14 Acme contracts. The find_contracts docstring said only "Find contracts." The team rewrites it: "Use first whenever the user names a company. Filters by counterparty, type and renewal month." They add 25 questions naming a company to the eval set. First-tool accuracy on that group goes from 40% to 96%, and the eval now fails if a later prompt edit breaks it.
Follow-up questions to expect
- "What is the first thing you look at?" — The first wrong step in the trace: which tool was chosen and with what arguments.
- "How do you debug in production without logging private data?" — Trace with LangSmith or OpenTelemetry, mask personal fields before export, and sample traces rather than keeping every one.
- "How do you know a fix did not break something else?" — Re-run the full eval set and compare tool accuracy and success rate per question type, not just the average.