LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you debug LangChain retrieval issues?


What you need to know

The first question: retrieval or generation?

  1. Find the right chunk yourself — search the source documents for the answer. Is it in the corpus at all?
  2. Is it in the index? — query the vector store for its id or source; the loader may have skipped the file.
  3. Did the retriever return it? — run retriever.invoke(question) and look.
  4. Did it reach the prompt? — open the model step in the trace and read the rendered context.
  5. Did the model use it? — if the chunk is there and the answer is wrong, fix the prompt or model.

Inspect the results

Python
question = "Can I carry forward unused earned leave?"docs = retriever.invoke(question)print(len(docs))for d in docs:    print(d.metadata.get("source"), d.metadata.get("page"), "|", d.page_content[:120])for doc, score in vector_store.similarity_search_with_score(question, k=10):    print(round(score, 3), doc.metadata.get("source"))

Look at the top 10, not just the top 4. If the right chunk sits at rank 7, raise k or add reranking. If it is nowhere, it is a chunking, embedding or ingestion problem.

Causes, roughly in order of frequency

SymptomCauseCheck
Zero documentsFilter matches nothing, or threshold too highRun without the filter and threshold
Right topic, wrong detailChunks too big — one vector averages several topicsLook at chunk lengths
Fragments without meaningChunks too small, or headings lostKeep headings in chunk text or metadata
Nonsense results for everythingDifferent embedding model for index and queryCompare the model recorded with the index
Misses exact codes or namesPure vector searchAdd BM25 / hybrid search
Old answersStale or duplicate chunksCheck the indexing job and record manager
Follow-up questions fail"What about next year?" searched as-isRewrite into a standalone query first

Make it repeatable

When you find a failing question, add it with its correct chunk id to your retrieval test set. The fix is done when recall@k on the whole set does not drop.

A real-life example

An HR bot answers "Can I carry forward unused earned leave?" with "No" — the handbook says yes, up to 30 days.

The engineer runs the retriever: the 4 results are about sick leave and casual leave. similarity_search_with_score with k=10 shows the carry-forward chunk at rank 9. Reading it reveals the problem: the chunk starts mid-table with "…30 days | Yes | Annual" and has no heading, because the PDF loader split the table from its title "Earned Leave". The embedding has almost no meaning to match.

They switch to a loader that keeps table headers, and prefix each chunk with its section heading. The chunk now ranks first. The question goes into the test set, and recall@5 on the 150-question set rises from 0.81 to 0.87.

Follow-up questions to expect

  • "How do you debug retrieval in an agent?" — Look at the retrieval tool's calls in the trace: the query the model wrote is often worse than the user's question; improve the tool description or rewrite queries.
  • "What if scores are all very close?" — The embedding is not separating your content well; check chunking, then try a domain-suited embedding model or reranking.
  • "How do you catch ingestion gaps?" — Count documents and chunks per source after each indexing run and alert on drops.