Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Retriever-Reader Pattern
A support team had 12,000 resolved tickets and a simple wish: let an agent ask a question in plain English and get the answer from the ticket history. Their first design was the obvious one. For each incoming question, loop over all 12,000 tickets, ask the language model "does this ticket answer the question, and if so what is the answer?", and collect whatever comes back.
Do the arithmetic before you judge them. The average ticket is about 800 tokens. Twelve thousand tickets is 9,600,000 input tokens per question. At 3 dollars per million input tokens that is 28.80 dollars to answer one question. Run the calls sequentially at 600 ms each and you wait two hours; parallelise fifty at a time and you still wait two and a half minutes, while your rate limit screams.
The design is not merely expensive. It is structurally wrong, and the fix reveals the pattern that sits underneath essentially every question-answering system built in the last decade. The insight: deciding that 11,995 of those tickets are irrelevant does not require a language model. It requires something cheap. Only the last handful of candidates deserves expensive, careful reading.
That split — a cheap fast filter followed by an expensive accurate reader — is the retriever-reader pattern.
The shape of the architecture
12,000 documents | ============ RETRIEVER ============== cheap, approximate, high recall vector search over pre-built index cost: ~0 dollars latency: ~30 ms | top 5 candidates | ============= READER ================ expensive, precise, high accuracy language model reads the candidates cost: ~0.012 dollars latency: ~1,300 ms | grounded answer + citationsNotice what each stage is optimised for, because they are deliberately different objectives:
| Retriever | Reader | |
|---|---|---|
| Job | Do not lose the right document | Extract the right answer |
| Optimises for | Recall — it is fine to return junk alongside | Precision — it must be right |
| Sees | All 12,000 documents | 5 documents |
| Cost per query | Fractions of a cent | Cents |
| Failure mode | Right document never surfaces | Right document surfaces, answer still wrong |
The retriever's job is not to be right. It is to make sure the right answer is somewhere in the small pile it hands to the reader. Precision is the reader's problem.
This asymmetry is the reason the pattern works. A retriever that returns 5 documents of which 1 is correct has done a perfect job — it reduced the search space by a factor of 2,400 without losing the answer. A retriever that returns 5 beautifully precise documents, none of which contains the answer, has failed completely, and there is nothing the reader can do about it.
Designing the retriever: four decisions
Decision 1 — Chunk size
You cannot index whole documents as single units. An embedding is one fixed-length vector; forcing a 6,000-word document into 768 numbers averages away everything specific. A document about billing, shipping, and returns produces a vector that is close to nothing in particular — a phenomenon worth naming: topic dilution.
So you split. How small?
| Chunk size | What it does well | What breaks | Suits |
|---|---|---|---|
| 100–200 tokens | Very precise matching; each chunk is one idea | Answers spanning two sentences get split; pronouns lose their referent | FAQ pairs, glossaries, product specs |
| 300–500 tokens | Balanced — a full paragraph or two of context | Occasionally mixes two topics | Most prose: docs, policies, articles |
| 800–1,200 tokens | Preserves long arguments and procedures | Topic dilution; wastes context budget on irrelevant text | Narrative documents, legal reasoning, tutorials |
Overlap exists to stop answers being guillotined at a boundary. If a chunk ends mid-explanation, neither the chunk before nor the chunk after contains the complete answer. Overlapping by 10–20% of the chunk size means every sentence appears in at least one chunk with its surrounding context.
The cost of overlap is real and worth computing. A 100,000-token corpus split into 500-token chunks with no overlap gives 200 chunks. With 100 tokens of overlap, each chunk advances only 400 tokens, so you get ceil(100000 / 400) = 250 chunks — 25% more vectors to store, embed and search. Push overlap to 250 (50%) and you get 400 chunks, double the index, for diminishing returns.
Overlap is insurance against boundary cuts, and like all insurance it has a premium. Ten to twenty per cent buys most of the protection; fifty per cent mostly buys you a bigger bill.
Splitting should respect structure, not character counts. Break on paragraph breaks first, then sentences, then words — never mid-word. That is exactly what a recursive splitter does.
Decision 2 — Embedding strategy
Two rules here that people break constantly.
The same model must embed both queries and chunks. Different models produce vectors in incompatible spaces. Comparing a vector from model A with a vector from model B gives numbers that look like similarities and mean nothing. If you change embedding model, you must re-embed the entire corpus. There is no migration path.
Many models are asymmetric and require prefixes. A question ("how do I reset my password?") and a passage ("Password resets are handled via the account settings page...") do not look alike lexically or structurally. Models such as the E5 and BGE families are trained with instruction prefixes to bridge this — "query: " for questions, "passage: " for documents. Forgetting the prefix silently drops retrieval quality by several points, with no error message. Check the model card; this is the single most common silent misconfiguration in RAG.
Decision 3 — Top-k
How many chunks to hand the reader. This is a direct recall-versus-noise trade, and it is measurable.
| k | Recall on a 200-question test set (illustrative run) | Context tokens (at 400/chunk) | Effect on the reader |
|---|---|---|---|
| 1 | 0.62 | 400 | Cheapest; fails outright 38% of the time |
| 3 | 0.81 | 1,200 | Good default for narrow, factual questions |
| 5 | 0.88 | 2,000 | The usual sweet spot |
| 10 | 0.93 | 4,000 | Diminishing returns begin; distractors accumulate |
| 20 | 0.95 | 8,000 | Answer quality often drops despite higher recall |
That last row is the counter-intuitive one and it catches people out. Going from k=10 to k=20 lifted recall by 2 points but doubled the context and buried the correct passage among nineteen distractors. Models degrade when relevant information sits in the middle of a long context — the "lost in the middle" effect — and every irrelevant chunk is a chance for the model to answer a slightly different question than the one asked.
The way out of this trade is not a bigger k. It is retrieving wide and then re-ranking down: fetch 50, score them properly, keep 5. You get the recall of k=50 with the context size of k=5.
Decision 4 — Similarity metric
Three metrics show up, and the relationship between them matters more than the choice.
Cosine similarity measures the angle between vectors, ignoring their lengths:
Dot product is the numerator alone — it rewards both alignment and magnitude. Euclidean distance measures straight-line separation.
The practical point: if all vectors are normalised to unit length, cosine and dot product give identical rankings, and Euclidean distance gives the same ordering too. Concretely, for unit vectors, squared Euclidean distance equals 2 − 2·cos. So a cosine of 0.9 corresponds to a squared distance of 2 − 1.8 = 0.2, and a cosine of 0.5 to 2 − 1.0 = 1.0. Higher cosine, lower distance, same order.
This is why the standard recipe is: normalise your vectors, then use inner product. It is the fastest operation, and it is exactly cosine similarity. The one time it matters is if you do not normalise and use raw dot product — then long vectors win regardless of relevance, and documents with unusual token distributions dominate every result list for no good reason.
Building the retriever
1import numpy as np, faiss2from sentence_transformers import SentenceTransformer34class Retriever:5 def __init__(self, model_name="BAAI/bge-small-en-v1.5"):6 self.model = SentenceTransformer(model_name)7 self.chunks = []8 self.index = None910 def _encode(self, texts, is_query=False):11 # BGE is asymmetric: queries get an instruction prefix, passages do not12 prefix = "Represent this sentence for searching relevant passages: "13 texts = [prefix + t for t in texts] if is_query else texts14 v = self.model.encode(texts, normalize_embeddings=True)15 return np.asarray(v, dtype="float32")1617 def build(self, chunks):18 self.chunks = chunks19 vecs = self._encode([c["text"] for c in chunks])20 self.index = faiss.IndexFlatIP(vecs.shape[1]) # inner product == cosine21 self.index.add(vecs)2223 def search(self, query, k=5, min_score=0.0, where=None):24 scores, ids = self.index.search(self._encode([query], is_query=True), k * 4)25 out = []26 for s, i in zip(scores[0], ids[0]):27 if i == -1 or s < min_score:28 continue29 c = self.chunks[i]30 if where and not all(c["meta"].get(f) == v for f, v in where.items()):31 continue32 out.append({**c, "score": float(s)})33 if len(out) == k:34 break35 return outTwo details worth copying. The retriever over-fetches (k * 4) so that score thresholds and metadata filters do not leave you with fewer results than requested. And filtering happens after the vector search here — fine for small corpora, but at scale you want the vector database to apply the filter during search, otherwise a narrow filter over millions of vectors returns nothing at all.
Measuring the retriever on its own
Evaluate the retriever separately from the reader. Otherwise you cannot tell which half is broken. You need a set of questions each labelled with the chunk id that answers it.
| Metric | Question it answers | Formula |
|---|---|---|
| Recall@k / Hit rate | Did the right chunk appear in the top k at all? | fraction of queries where a relevant chunk is in the top k |
| MRR | How near the top was the first right chunk? | mean of 1/rank of first relevant hit |
| Precision@k | How much of what we returned was useful? | relevant hits in top k, divided by k |
A worked MRR. Five test questions; the rank at which the correct chunk first appeared was 1, 3, 2, not-found, 1. The reciprocal ranks are 1.000, 0.333, 0.500, 0.000, 1.000. Their sum is 2.833, divided by 5 gives MRR = 0.567. Recall@5 for the same run is 4/5 = 0.80.
Read those two numbers together. Recall 0.80 says the retriever usually finds the answer. MRR 0.567 says it often buries it at rank 2 or 3. That combination is the exact signature of a system that needs a re-ranker rather than a better embedding model — the candidates are there, the ordering is wrong.
Designing the reader
The reader takes the question and the surviving chunks and produces an answer. There are two families, and they are genuinely different tools.
Extractive readers: point at a span
An extractive reader selects a contiguous span of the source text as the answer. Classic BERT-style QA models do this by predicting two probability distributions over token positions — where the answer starts and where it ends — and taking the span that maximises the combined score.
1import torch2from transformers import AutoTokenizer, AutoModelForQuestionAnswering34# transformers 5 removed the "question-answering" pipeline; call the model directly5name = "deepset/roberta-base-squad2"6tok = AutoTokenizer.from_pretrained(name)7model = AutoModelForQuestionAnswering.from_pretrained(name)89def extract(question, context):10 enc = tok(question, context, return_tensors="pt", return_offsets_mapping=True)11 offsets = enc.pop("offset_mapping")[0]12 with torch.no_grad():13 out = model(**enc)14 in_ctx = torch.tensor([s == 1 for s in enc.sequence_ids(0)]) # context tokens only15 p_start = out.start_logits[0].masked_fill(~in_ctx, -1e9).softmax(-1)16 p_end = out.end_logits[0].masked_fill(~in_ctx, -1e9).softmax(-1)17 s = int(p_start.argmax())18 e = s + int(p_end[s:].argmax()) # the end cannot come before the start19 a, b = int(offsets[s][0]), int(offsets[e][1])20 return {"answer": context[a:b], "score": round(float(p_start[s] * p_end[e]), 2),21 "start": a, "end": b}2223ctx = ("Enterprise annual contracts carry a 14-day refund window "24 "measured from the invoice date. Monthly plans carry 30 days.")25print(extract("What is the refund window for enterprise annual contracts?", ctx))26# {'answer': '14-day', 'score': 0.87, 'start': 36, 'end': 42}The span is guaranteed to exist verbatim in the source, so fabrication is impossible by construction. The span score is a real probability from the model, and a low one is a useful signal to abstain — though, like any model score, its threshold has to be tuned on your own data. It runs in about 20 ms on CPU. But it cannot synthesise across two passages, cannot rephrase, cannot say "there are three conditions and they are...", and it fails whenever the answer is not a literal substring.
Generative readers: write the answer
A generative reader is a language model given the context and told to answer from it. It can combine facts from several chunks, resolve pronouns, reformat, summarise, and handle conversational phrasing. It is what almost everyone means by RAG today. The price is that it can fabricate, its confidence is not calibrated, it costs cents rather than nothing, and it takes over a second.
| Extractive | Generative | |
|---|---|---|
| Output | Exact span from the source | Freely written text |
| Can hallucinate | No, structurally impossible | Yes, must be constrained by prompt |
| Multi-passage synthesis | No | Yes |
| Confidence score | A span probability you can threshold | None you can trust |
| Latency / cost | ~20 ms, free after hosting | ~1,300 ms, per-token pricing |
| Choose when | Answers are literal values; audit trail is paramount; huge query volume | Answers need explanation, synthesis or conversation |
Reader design decisions
Keep sampling randomness low. Any randomness in a grounded-answering task is pure downside — you get variability with no creativity benefit, and you lose reproducibility when debugging. Set temperature to 0 where the model accepts it. Some current reasoning models reject a temperature setting or only allow the default; with those, pin the exact model version and let the prompt contract below do the constraining.
Number the chunks and demand citations. Citations are not a nicety; they are the mechanism by which a human can falsify the answer. Numbering also lets you programmatically check that every cited index actually existed.
Order chunks deliberately. Given the lost-in-the-middle effect, one effective trick is to place the highest-scoring chunk first and the second-highest last, pushing weak candidates into the middle where the model attends least.
Make abstention easy and specific. A vague instruction to "say if you don't know" is usually ignored. An exact required string is not.
1READER_PROMPT = """You answer strictly from the numbered passages below.23Rules:41. Every factual claim ends with its passage number, e.g. [2].52. If the passages do not contain the answer, reply with exactly:6 NOT_IN_CONTEXT73. If two passages conflict, report both and cite each.84. Do not add information from your own knowledge.910Passages:11{context}1213Question: {question}14Answer:"""1516def read(client, question, chunks):17 # strongest first, second-strongest last, weak ones buried in the middle18 ordered = chunks[:1] + chunks[2:] + chunks[1:2] if len(chunks) > 2 else chunks19 ctx = "\n\n".join(20 f"[{i+1}] (source: {c['meta']['source']})\n{c['text']}"21 for i, c in enumerate(ordered))22 resp = client.chat.completions.create(23 model=CHAT_MODEL, # from config, as in the first lesson24 messages=[{"role": "user",25 "content": READER_PROMPT.format(context=ctx, question=question)}])26 text = resp.choices[0].message.content.strip()27 if text == "NOT_IN_CONTEXT":28 return {"answer": None, "sources": []}29 used = sorted({int(n) for n in __import__("re").findall(r"\[(\d+)\]", text)})30 return {"answer": text,31 "sources": [ordered[n-1]["meta"]["source"] for n in used32 if 1 <= n <= len(ordered)]}Where this pattern goes wrong
| Symptom | Actual cause | Fix |
|---|---|---|
| Answer is confidently wrong and cites nothing | Reader ignored the grounding instruction; no abstention path | Exact abstention string; reject uncited claims in post-processing |
| "I don't have that information" on questions you know are covered | Retriever failure — right chunk never surfaced | Measure Recall@k; check embedding prefixes; add a re-ranker |
| Answers are half-sentences or miss a condition | Chunks too small; answer split across a boundary | Increase chunk size or overlap; use parent-document expansion |
| Retrieval scores all cluster around 0.7–0.8, right and wrong alike | Topic dilution from oversized chunks | Shrink chunks; split on structure not character count |
| Quality collapsed overnight with no code change | Embedding model version changed; index and queries now in different spaces | Pin the model version; re-embed the whole corpus on any change |
| Works in testing, wrong in production on recent topics | Index is stale relative to the documents | Incremental re-indexing on document change events |
The two entries worth dwelling on are the mirror images of each other. A retriever failure looks like a reader failure and vice versa, and teams routinely spend a week tuning prompts when the correct chunk was never in the context at all. The diagnostic takes thirty seconds: log the retrieved chunks for a failing query and read them yourself. If the answer is not in them, stop touching the prompt.
Before you tune the reader, read what the retriever gave it. Most "the model is hallucinating" bug reports are "the model was handed the wrong five paragraphs" bug reports.
Wiring both halves together
1class RetrieverReaderQA:2 def __init__(self, retriever, client, k=5, min_score=0.30):3 self.retriever, self.client = retriever, client4 self.k, self.min_score = k, min_score56 def ask(self, question, where=None):7 hits = self.retriever.search(question, k=self.k,8 min_score=self.min_score, where=where)9 if not hits:10 return {"answer": None, "reason": "no_candidates", "sources": []}1112 # a weak best hit is a strong signal that the corpus does not cover this13 if hits[0]["score"] < 0.45:14 return {"answer": None, "reason": "low_confidence",15 "top_score": hits[0]["score"], "sources": []}1617 result = read(self.client, question, hits)18 result["retrieval_scores"] = [round(h["score"], 3) for h in hits]19 return resultThe low_confidence branch is the cheapest quality win available. If the best chunk in a 61,000-chunk index only scores 0.42 against the query, the corpus almost certainly does not cover the question, and calling the reader will cost you money to produce a plausible fabrication. Refusing before generation is faster, cheaper and more honest. Calibrate the threshold on your own data — score distributions differ sharply between embedding models, so a number copied from a blog post will be wrong for you.
Making it better once it works
Two-stage retrieval
Fetch 50 candidates with the fast bi-encoder, score all 50 with a cross-encoder that reads query and passage together, keep the top 5. It costs 50–150 ms and is the direct fix for the "high recall, low MRR" signature in the worked example above: the right chunk is already in the pile, and the cross-encoder moves it to the top. This is the highest-value single change in most systems.
Parent-document retrieval
Index small chunks (200 tokens) for precise matching, but when a chunk is retrieved, give the reader its parent — the 1,000-token section it came from. You match precisely and read with full context, which resolves the classic "chunk was too small to contain the whole condition" failure without inflating the index.
Sentence-window retrieval
The same idea at finer granularity: index individual sentences, but expand each retrieved sentence to include the three sentences either side before handing it to the reader.
Contextual retrieval
A chunk that says "The window is 14 days" has lost the fact that it came from the enterprise annual contract, not the monthly plan. Contextual retrieval, a technique Anthropic published in September 2024, repairs this at indexing time: for each chunk, a model reads the whole document and writes a short note (about 50–100 tokens) situating the chunk — "From the 2025 enterprise terms, section 9, on refunds for annual contracts". That note is prepended to the chunk before it is embedded and before it goes into the keyword index. In Anthropic's own tests, this cut top-20 retrieval failures by 35% with contextual embeddings alone, 49% when combined with contextual BM25, and 67% with re-ranking added on top. The cost is one model call per chunk at indexing time, which prompt caching of the shared document makes much cheaper.
Metadata filtering
Restrict the search space before similarity is computed — department, date range, document status, the requesting user's access level. Filtering out superseded documents is not an optimisation, it is a correctness requirement: a retired policy reads exactly like a current one and will score just as highly.
Caching
Support questions repeat heavily. Cache the embedding of every query (deterministic, so caching is free correctness-wise) and cache full answers for exact repeats with a short expiry.
What this means when you build one
Design the two halves as separate services with separate tests, even if they run in the same process. The retriever's contract is "given a question, return candidate passages with scores"; the reader's contract is "given a question and passages, return a grounded answer or abstain". Keep that boundary clean and every future decision gets easier: you can swap the embedding model without touching the prompt, add a re-ranker as a middle stage, or replace the generative reader with an extractive one for a high-volume endpoint.
Instrument the boundary. Log, for every query, the retrieved chunk ids and scores alongside the final answer. This one log line turns debugging from guesswork into a lookup — when a user reports a bad answer you can see immediately whether the right passage was even present. Teams that skip this end up arguing about prompts for weeks.
Finally, set your thresholds from measurement, not intuition. Build a set of 50 real questions with the correct chunk labelled for each — an afternoon's work. Run the retriever alone and compute Recall@5 and MRR. Recall tells you the ceiling on your whole system's accuracy; no reader can exceed it. MRR tells you whether you need better embeddings (low recall and low MRR) or just better ordering (high recall, low MRR). Those two numbers will direct your next month of work more reliably than any amount of reading individual outputs and forming impressions.