Course Content
LangChain Mastery
7 sections · 109 lessons
How do you debug a LangChain chain?
What you need to know
A chain fails in one of a few places: the input shape, the rendered prompt, the model's reply, the parser, or a custom function. Debugging means finding which one.
- Trace it — turn on LangSmith and look at the failing run's step tree.
- Read the rendered prompt — most "bad answers" come from a prompt that did not contain what you thought.
- Run prefixes —
prompt.invoke(x), then(prompt | llm).invoke(x), then the full chain. - Tap values — insert a step that logs and returns its input.
- Fix and add a test — keep the failing input as a regression case.
Tools
1from langchain_core.globals import set_debug2set_debug(True) # logs every step's inputs and outputs34print(prompt.invoke(inputs).to_string()) # see the exact prompt5print((prompt | llm).invoke(inputs)) # raw AIMessage, before parsing67def tap(x):8 print("between steps:", repr(x)[:500])9 return x10chain = prompt | llm | tap | parser # a function becomes a RunnableLambdaset_debug(True)prints everything, including full prompts; use it locally, not in production logs.set_verbose(True)prints less.chain.get_graph().print_ascii()draws the chain's structure (it needs thegrandalfpackage installed).async for ev in chain.astream_events(inputs, version="v2")yields start, stream and end events for every step, useful when a streaming UI stalls.
Common causes by symptom
| Symptom | Usual cause |
|---|---|
KeyError: missing variables | Input dict key does not match the template |
| Answer ignores the documents | Retriever returned nothing, or the context variable was empty |
OutputParserException | Model added text around the JSON; use structured output |
content='...' stored in the DB | Missing StrOutputParser |
| Works locally, fails in prod | Different model, settings or package version |
A real-life example
A support bot over an ed-tech company's help-centre docs started answering "I don't know" to 30% of questions after a release. Unit tests passed. In LangSmith, the rendered prompt for a failing run showed Context: followed by nothing. Running the retrieval step alone returned zero documents, and the metadata filter had changed from product="app" to product="App" in the new code. The fix was one line, and the team added the failing question to their test set with a check that retrieval returns at least one document. Total debugging time was about 20 minutes because they looked at the trace first instead of guessing at the prompt.
Follow-up questions to expect
- "How do you debug without LangSmith?" —
set_debug(True), a callback that logs inputs and outputs, and running each prefix of the chain by hand. - "How do you find which runs to look at?" — Tag runs with
run_name,tagsandmetadata(user, prompt version) and filter on them. - "How do you debug a problem you can't reproduce?" — Pull the exact inputs from the production trace and replay them locally with the same model settings.