Course Content
RAG Systems
12 sections · 66 lessons
What is the specific role of the Retrieval component?
What you need to know
Mechanically, a retriever does four things: it turns the query into a search form (an embedding, keywords or both), searches the index, applies filters such as tenant and version, and returns ranked passages with text, metadata and scores.
Why it sets the ceiling
Think of the model as a very good reader who can only read the pages you hand them. If the page with the answer is not in the pile, a better reader does not help. That is why a retrieval miss often shows up as a hallucination: the model answers from memory instead.
Measuring it: a tiny worked example
Five labelled HR questions. For each, the rank at which the correct chunk came back:
| Question | Rank of correct chunk | In top 3? | Reciprocal rank |
|---|---|---|---|
| Paternity leave days | 1 | yes | 1.00 |
| Carry-forward limit | 3 | yes | 0.33 |
| Notice period band 4 | not in top 10 | no | 0 |
| Notice buy-out rule | 2 | yes | 0.50 |
| Earned leave per year | 1 | yes | 1.00 |
recall@1 = 2 / 5 = 0.40recall@3 = 4 / 5 = 0.80MRR = (1 + 0.33 + 0 + 0.5 + 1) / 5 = 0.57The one miss (notice period band 4) tells you more than the averages: go and read why that chunk is not found. The full set of retrieval metrics, including nDCG, is covered in the evaluation section.
Building the labelled set
Take 50 to 100 real questions from logs or from subject experts. For each, record which chunk (or document and section) holds the answer. This takes a day and is the most valuable asset in a RAG project, because every change after that can be measured.
A real-life example
A legal-contract search tool answers "What is the governing law of the contract with Vendor A?" wrongly: it says "Delaware". The contract says "the laws of India, with courts at Mumbai".
The engineer prints the top 5 retrieved chunks. None is from Vendor A's contract. They are governing-law clauses from other contracts, which are almost identical in wording. The retriever found the right kind of clause from the wrong document.
The fix is in retrieval, not in generation: extract the counterparty name as metadata at ingest, and filter by it when the question names a vendor. Recall@5 on their 80-question test set rises from 0.71 to 0.93 in their measurements. The prompt did not change.
Follow-up questions to expect
- "What k do you evaluate at?" — At the k you actually pass to the model, and at the wider k you pass to a reranker. Recall@30 tells you whether a reranker could help; recall@5 tells you what the model sees.
- "How do you get labels without experts?" — Generate candidate questions from chunks with an LLM, then have a person check a sample. Real user questions are still better.
- "What is the difference between retrieval and a search engine?" — Very little. Retrieval is search tuned for a model reader: fewer, more precise passages, with metadata for citations.