LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug a LangChain chain that fails?


Bisect the chain: the first wrong output is the bugprompt.invoke — does it render?prompt plus llm — what did it reply?add the parser — does it accept it?retriever.invoke — any documents at all?
Every piece of a chain is a runnable you can call alone, so a trace or four calls find the failing step in minutes.

What you need to know

A chain is several steps glued together, so the stack trace often points at LangChain internals, not at your mistake. The job is to find which step failed and what it received.

Step 1: turn on tracing

Bash
export LANGSMITH_TRACING=trueexport LANGSMITH_API_KEY=...export LANGSMITH_PROJECT=support-bot-dev

No code change is needed. Every invoke now appears in LangSmith as a tree: prompt, model, parser, retriever, tool. The failed step is red and shows the exception and the exact inputs.

Step 2: debug output locally

Python
from langchain_core.globals import set_debugset_debug(True)       # prints every step's inputs and outputs with timings

In LangChain 1.x this lives in langchain_core.globals; the old langchain.globals module was removed.

Step 3: bisect the chain

Python
prompt.invoke({"question": q})                       # does the template render?(prompt | llm).invoke({"question": q})               # what does the model return?(prompt | llm | parser).invoke({"question": q})      # does the parser accept it?retriever.invoke(q)                                  # any documents at all?

Each piece is a runnable, so you can call it alone. The first call that looks wrong is your bug.

The four common causes

SymptomLikely cause
KeyError / "missing variables"Input dict key does not match the {placeholder} in the prompt
OutputParserExceptionModel reply is not in the expected format
Fluent but wrong answerRetriever returned zero or wrong documents
RateLimitError, APITimeoutError, 5xxProvider problem; needs retry, timeout or fallback

Many LangChain exceptions include a link to a troubleshooting page, such as OUTPUT_PARSING_FAILURE — read it.

A real-life example

A support bot's summary chain starts failing after a deploy with OutputParserException. The stack trace ends deep inside the JSON parser.

The engineer opens the failing LangSmith trace. The model step's output reads: "Sure! Here is the summary in JSON: ``json {...}`". A prompt edit that morning removed the line "Return only JSON". The parser was right to reject it. The fix was to switch the chain to llm.with_structured_output(Summary)`, which uses the provider's structured-output mode instead of relying on the prompt. Total time to find the cause: about five minutes, most of it opening the trace.

Follow-up questions to expect

  • "How do you debug without sending data to LangSmith?" — Use set_debug(True) locally, a custom callback handler that writes to your own logs, or a self-hosted LangSmith deployment.
  • "How do you reproduce a production failure?" — Take the inputs from the trace, add them to a dataset, and re-run that example locally.
  • "How do you debug an async or streaming chain?" — Same traces; plus astream_events to watch start and end events for each step.