RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How does LangSmith help in debugging RAG applications?


Seconds per span in one slow trace0.40.10.053.91.701234vector searchreranker,200 candidatesSpans: rewrite, embed, search, rerank, generate. Setting k back to 40 fixed it.
The trace pointed at a config change in the reranker, not at the vector search the team had planned to optimise.

What you need to know

Debugging a RAG answer is like debugging any pipeline: find the first stage with bad output. LangSmith shows every stage in one screen.

  1. Find the trace — search by user id, session id or time from the bug report.
  2. Retriever span — is the answer-bearing text in the returned chunks? Check scores and metadata too (right product, right version, right tenant).
  3. Rendered prompt — is the chunk actually in the final prompt, complete, and near the top?
  4. LLM span — given a good prompt, did the model use it? Did it cite? Did it add outside facts?
  5. Latency waterfall — which span took the time?
  6. Save and test — add the trace to a dataset, fix, and run an experiment to confirm the fix without breaking other cases.

What the first wrong step tells you

First wrong stepLikely causeTypical fix
Retriever: answer not in chunksMissing document, bad chunk split, wrong filter, keyword queryFix ingest, hybrid search, fix filter
Retriever: answer at rank 12, only top 4 usedWeak rankingAdd or tune reranker
Prompt: chunk cut offToken limit, bad truncationBudget by tokens, fewer chunks
LLM: ignores contextWeak instructions, conflicting chunksStronger grounding prompt, labelled sources
LLM: correct but slowLong prompt, big modelFewer chunks, smaller model, streaming

Turning a trace into a test

Python
from langsmith import Clientclient = Client()def retrieved_expected_doc(outputs: dict, reference_outputs: dict) -> bool:    return reference_outputs["doc_id"] in outputs["doc_ids"]client.evaluate(    lambda inputs: answer_question(inputs["question"]),   # your pipeline    data="faq-bot-regressions",                          # dataset of saved traces    evaluators=[retrieved_expected_doc],    experiment_prefix="reranker-v2",)

Each saved example has the question as input and the expected document id as reference output. The evaluator returns True if the pipeline retrieved it. experiment_prefix names the run so you can compare "reranker-v1" and "reranker-v2" side by side. I ran this pattern locally with an in-memory example and it scored as expected.

A real-life example

A bank's FAQ bot takes 6.2 seconds for some answers. The latency waterfall in one slow trace shows: query rewrite 0.4 s, embedding 0.1 s, vector search 0.05 s, reranker 3.9 s, generation 1.7 s. The vector search the team planned to optimise is under 1% of the time.

The reranker span shows it received 200 candidates, because a config change set k=200. Setting it back to 40 brings the reranker to under a second. The engineer adds a trace-based alert on reranker span duration and saves the trace to a "latency regressions" dataset.

Follow-up questions to expect

  • "How do you debug one tenant's bad answers?" — Filter traces by tenant_id and look at retrieval scores and empty results. Tenant-specific failures are usually missing documents or wrong filters.
  • "How do you compare two chunking strategies?" — Build two indexes, run both over the same dataset as two experiments, and compare recall and answer scores in the comparison view.
  • "How do you connect user feedback to traces?" — Send thumbs up or down with the run id through the feedback API, then filter traces by negative feedback.