LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug a LangChain chain?


What you need to know

A chain fails in one of a few places: the input shape, the rendered prompt, the model's reply, the parser, or a custom function. Debugging means finding which one.

  1. Trace it — turn on LangSmith and look at the failing run's step tree.
  2. Read the rendered prompt — most "bad answers" come from a prompt that did not contain what you thought.
  3. Run prefixes — prompt.invoke(x), then (prompt | llm).invoke(x), then the full chain.
  4. Tap values — insert a step that logs and returns its input.
  5. Fix and add a test — keep the failing input as a regression case.

Tools

Python
from langchain_core.globals import set_debugset_debug(True)                      # logs every step's inputs and outputsprint(prompt.invoke(inputs).to_string())       # see the exact promptprint((prompt | llm).invoke(inputs))           # raw AIMessage, before parsingdef tap(x):    print("between steps:", repr(x)[:500])    return xchain = prompt | llm | tap | parser            # a function becomes a RunnableLambda
  • set_debug(True) prints everything, including full prompts; use it locally, not in production logs.
  • set_verbose(True) prints less.
  • chain.get_graph().print_ascii() draws the chain's structure (it needs the grandalf package installed).
  • async for ev in chain.astream_events(inputs, version="v2") yields start, stream and end events for every step, useful when a streaming UI stalls.

Common causes by symptom

SymptomUsual cause
KeyError: missing variablesInput dict key does not match the template
Answer ignores the documentsRetriever returned nothing, or the context variable was empty
OutputParserExceptionModel added text around the JSON; use structured output
content='...' stored in the DBMissing StrOutputParser
Works locally, fails in prodDifferent model, settings or package version

A real-life example

A support bot over an ed-tech company's help-centre docs started answering "I don't know" to 30% of questions after a release. Unit tests passed. In LangSmith, the rendered prompt for a failing run showed Context: followed by nothing. Running the retrieval step alone returned zero documents, and the metadata filter had changed from product="app" to product="App" in the new code. The fix was one line, and the team added the failing question to their test set with a check that retrieval returns at least one document. Total debugging time was about 20 minutes because they looked at the trace first instead of guessing at the prompt.

Follow-up questions to expect

  • "How do you debug without LangSmith?" — set_debug(True), a callback that logs inputs and outputs, and running each prefix of the chain by hand.
  • "How do you find which runs to look at?" — Tag runs with run_name, tags and metadata (user, prompt version) and filter on them.
  • "How do you debug a problem you can't reproduce?" — Pull the exact inputs from the production trace and replay them locally with the same model settings.