Course Content
RAG Systems
12 sections · 66 lessons
How does LangSmith help in debugging RAG applications?
What you need to know
Debugging a RAG answer is like debugging any pipeline: find the first stage with bad output. LangSmith shows every stage in one screen.
- Find the trace — search by user id, session id or time from the bug report.
- Retriever span — is the answer-bearing text in the returned chunks? Check scores and metadata too (right product, right version, right tenant).
- Rendered prompt — is the chunk actually in the final prompt, complete, and near the top?
- LLM span — given a good prompt, did the model use it? Did it cite? Did it add outside facts?
- Latency waterfall — which span took the time?
- Save and test — add the trace to a dataset, fix, and run an experiment to confirm the fix without breaking other cases.
What the first wrong step tells you
| First wrong step | Likely cause | Typical fix |
|---|---|---|
| Retriever: answer not in chunks | Missing document, bad chunk split, wrong filter, keyword query | Fix ingest, hybrid search, fix filter |
| Retriever: answer at rank 12, only top 4 used | Weak ranking | Add or tune reranker |
| Prompt: chunk cut off | Token limit, bad truncation | Budget by tokens, fewer chunks |
| LLM: ignores context | Weak instructions, conflicting chunks | Stronger grounding prompt, labelled sources |
| LLM: correct but slow | Long prompt, big model | Fewer chunks, smaller model, streaming |
Turning a trace into a test
1from langsmith import Client23client = Client()45def retrieved_expected_doc(outputs: dict, reference_outputs: dict) -> bool:6 return reference_outputs["doc_id"] in outputs["doc_ids"]78client.evaluate(9 lambda inputs: answer_question(inputs["question"]), # your pipeline10 data="faq-bot-regressions", # dataset of saved traces11 evaluators=[retrieved_expected_doc],12 experiment_prefix="reranker-v2",13)Each saved example has the question as input and the expected document id as reference output. The evaluator returns True if the pipeline retrieved it. experiment_prefix names the run so you can compare "reranker-v1" and "reranker-v2" side by side. I ran this pattern locally with an in-memory example and it scored as expected.
A real-life example
A bank's FAQ bot takes 6.2 seconds for some answers. The latency waterfall in one slow trace shows: query rewrite 0.4 s, embedding 0.1 s, vector search 0.05 s, reranker 3.9 s, generation 1.7 s. The vector search the team planned to optimise is under 1% of the time.
The reranker span shows it received 200 candidates, because a config change set k=200. Setting it back to 40 brings the reranker to under a second. The engineer adds a trace-based alert on reranker span duration and saves the trace to a "latency regressions" dataset.
Follow-up questions to expect
- "How do you debug one tenant's bad answers?" — Filter traces by
tenant_idand look at retrieval scores and empty results. Tenant-specific failures are usually missing documents or wrong filters. - "How do you compare two chunking strategies?" — Build two indexes, run both over the same dataset as two experiments, and compare recall and answer scores in the comparison view.
- "How do you connect user feedback to traces?" — Send thumbs up or down with the run id through the feedback API, then filter traces by negative feedback.