Course Content
RAG Systems
12 sections · 66 lessons
Why should you inspect the documents returned by the retriever?
What you need to know
The one-minute inspection
1query = "What notice period do I have to serve?"2for doc, dist in store.similarity_search_with_score(query, k=5):3 print(f"{dist:.3f}", doc.metadata.get("source"), doc.metadata.get("section"),4 "|", doc.page_content[:90].replace("\n", " "))On the HR store used earlier in this section, the first two lines were:
0.158 hr/handbook-v5.md Notice period | ## Notice period All employees must serve a notice period of 30 days.0.189 hr/handbook-v7.md Notice period | ## Notice period Employees in bands 1 to 3 must serve a notice period...In ten seconds you can see the bug: a superseded handbook ranked first. No amount of prompt work would have found that.
What to check
| Look at | Healthy | Warning sign |
|---|---|---|
| Chunk text | A coherent passage with its heading | Starts mid-sentence, orphan heading, table fragment |
| Scores | Clear gap between top hits and the rest | All scores flat and low: nothing really matched |
| Metadata | Right source, current version, right tenant | Old version, other country, draft document |
| Duplicates | Distinct passages | The same paragraph in several slots |
| Score type | You know if it is distance or similarity | Threshold set in the wrong direction |
Decide which half failed
- Is the answer in the retrieved chunks? If no, it is a retrieval problem. Go to step 2. If yes, it is a generation problem: check prompt, order and model.
- Is the answer in the index at all? Search for a key phrase directly. If it is missing, the loader or parser dropped it.
- Is it in the index but ranked low? Then try hybrid search, a reranker, a better chunk boundary or a filter.
Make it permanent
For every production request, log the query, the rewritten query, retrieved chunk IDs, scores, the prompt version and the answer. Tracing tools such as LangSmith, Langfuse or Arize Phoenix show this per request. Over time, questions with thumbs-down feedback plus their logged chunks become labelled examples for your evaluation set.
A real-life example
A bank's product-FAQ bot tells a customer the savings account minimum balance is Rs 5,000; the current figure is Rs 10,000. The first reaction is to change the prompt to "be careful with numbers".
An engineer instead replays the request from the trace. The top chunk is from a PDF of a different account variant, a basic savings account, whose title was lost during parsing, so its chunk just says "Minimum balance: Rs 5,000". The model read the right number from the wrong document. The fix is at ingestion: keep each PDF's product name in metadata and prepend it to every chunk, then filter by product when the question names one. The trace took five minutes to read; prompt changes would never have fixed it.
Follow-up questions to expect
- "How do you inspect at scale?" — Sample logged requests weekly, especially low-rated ones, and compute automatic metrics such as context relevance on the sample.
- "What if all scores are low?" — Treat it as "nothing found": answer that the documents do not cover the question, and log it as a content gap.
- "Is inspecting chunks a privacy risk?" — It can be. Restrict who can view traces, and mask personal data in logs.