LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug LangChain agent decision-making?


The trace of one wrong answerAcmerenewals?semantic_search5 wrongchunksnonefound0123firstwrong stepconfident,wrongfind_contracts would have returned all 14 Acme contracts; its docstring only said Find contracts.
Find the first step that went wrong — here the tool choice — and the fix is usually a description, not a bigger model.

What you need to know

Classify the failure first

When an agent gives a wrong answer, find which step went wrong. There are only a few possibilities:

Symptom in the traceLikely cause
Called the wrong toolOverlapping or vague descriptions
Right tool, wrong argumentsMissing field descriptions or examples in the schema
Right result, wrong answerTool output too long or unclear; model misread it
No tool call, guessed the answerSystem prompt does not require tools for facts
Same call repeated many timesError message gives no guidance; no call limit

Seeing the steps locally

Python
result = agent.invoke({"messages": [{"role": "user", "content": question}]})for m in result["messages"]:    if m.type == "ai" and m.tool_calls:        for c in m.tool_calls:            print("CALL  ", c["name"], c["args"])    elif m.type == "tool":        print("RESULT", m.name, str(m.content)[:200])    elif m.type == "ai":        print("ANSWER", m.content[:200])

The message list is the whole story of the run. For live output while it runs, use for chunk in agent.stream(inputs, stream_mode="updates"): print(chunk), which prints each node's update (the model's decision, then the tool results) as it happens. set_debug(True) from langchain_core.globals prints everything, but the volume makes it hard to read.

Tracing with LangSmith

Bash
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=support-bot-dev

With these set, every run is recorded with no code change. The trace shows what a print statement cannot: the full prompt including the tool schemas, token counts and latency per step, and which middleware changed the request. You can open a bad run, edit the prompt in the playground, and re-run that exact step.

Fixing, in order of leverage

  1. Tool descriptions — add "use for X, do not use for Y", and an example argument. This fixes most wrong choices.
  2. Tool count — merge or remove overlapping tools, or filter tools per request with middleware.
  3. Tool output — return a short, labelled result instead of raw JSON.
  4. System prompt — state the policy: "Always check the order tool before answering about delivery."
  5. Model — only then try a stronger model.

Make it a test

Build 30 to 100 real questions with the expected tool sequence, such as ["get_order", "track_shipment"]. Run them on every prompt, tool or model change and assert on the tools called, not only on the final text. LangSmith datasets and evaluators can run this; a plain pytest loop over result["messages"] also works.

A real-life example

A law firm's contract-search assistant keeps answering "Which contracts with Acme auto-renew?" with "I found no such contracts" — but they exist.

The trace shows the agent calling semantic_search("Acme auto-renew"), getting 5 chunks from unrelated contracts, and giving up. It never called find_contracts(counterparty="Acme"), which would have found all 14 Acme contracts. The find_contracts docstring said only "Find contracts." The team rewrites it: "Use first whenever the user names a company. Filters by counterparty, type and renewal month." They add 25 questions naming a company to the eval set. First-tool accuracy on that group goes from 40% to 96%, and the eval now fails if a later prompt edit breaks it.

Follow-up questions to expect

  • "What is the first thing you look at?" — The first wrong step in the trace: which tool was chosen and with what arguments.
  • "How do you debug in production without logging private data?" — Trace with LangSmith or OpenTelemetry, mask personal fields before export, and sample traces rather than keeping every one.
  • "How do you know a fix did not break something else?" — Re-run the full eval set and compare tool accuracy and success rate per question type, not just the average.