Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Retriever-Reader Pattern


A support team had 12,000 resolved tickets and a simple wish: let an agent ask a question in plain English and get the answer from the ticket history. Their first design was the obvious one. For each incoming question, loop over all 12,000 tickets, ask the language model "does this ticket answer the question, and if so what is the answer?", and collect whatever comes back.

Do the arithmetic before you judge them. The average ticket is about 800 tokens. Twelve thousand tickets is 9,600,000 input tokens per question. At 3 dollars per million input tokens that is 28.80 dollars to answer one question. Run the calls sequentially at 600 ms each and you wait two hours; parallelise fifty at a time and you still wait two and a half minutes, while your rate limit screams.

The design is not merely expensive. It is structurally wrong, and the fix reveals the pattern that sits underneath essentially every question-answering system built in the last decade. The insight: deciding that 11,995 of those tickets are irrelevant does not require a language model. It requires something cheap. Only the last handful of candidates deserves expensive, careful reading.

That split — a cheap fast filter followed by an expensive accurate reader — is the retriever-reader pattern.

Why the reader never sees 12,000 ticketsQuestion inplain EnglishRetrieverscores12,000 chunksTop 5 survive,in millisecondsReader getsonly those 5Answer, withthe ticketit came fromThe naive design asked the model about every ticket: 12,000 calls and about 40 dollars per question.
Retrieval is cheap and approximate, reading is expensive and precise — the pattern exists to put the expensive step behind the cheap one.

The shape of the architecture

Text
                    12,000 documents                           |   ============ RETRIEVER ==============   cheap, approximate, high recall   vector search over pre-built index   cost: ~0 dollars   latency: ~30 ms                           |                    top 5 candidates                           |   ============= READER ================   expensive, precise, high accuracy   language model reads the candidates   cost: ~0.012 dollars  latency: ~1,300 ms                           |                   grounded answer + citations

Notice what each stage is optimised for, because they are deliberately different objectives:

RetrieverReader
JobDo not lose the right documentExtract the right answer
Optimises forRecall — it is fine to return junk alongsidePrecision — it must be right
SeesAll 12,000 documents5 documents
Cost per queryFractions of a centCents
Failure modeRight document never surfacesRight document surfaces, answer still wrong

The retriever's job is not to be right. It is to make sure the right answer is somewhere in the small pile it hands to the reader. Precision is the reader's problem.

This asymmetry is the reason the pattern works. A retriever that returns 5 documents of which 1 is correct has done a perfect job — it reduced the search space by a factor of 2,400 without losing the answer. A retriever that returns 5 beautifully precise documents, none of which contains the answer, has failed completely, and there is nothing the reader can do about it.

Designing the retriever: four decisions

Decision 1 — Chunk size

You cannot index whole documents as single units. An embedding is one fixed-length vector; forcing a 6,000-word document into 768 numbers averages away everything specific. A document about billing, shipping, and returns produces a vector that is close to nothing in particular — a phenomenon worth naming: topic dilution.

So you split. How small?

Chunk sizeWhat it does wellWhat breaksSuits
100–200 tokensVery precise matching; each chunk is one ideaAnswers spanning two sentences get split; pronouns lose their referentFAQ pairs, glossaries, product specs
300–500 tokensBalanced — a full paragraph or two of contextOccasionally mixes two topicsMost prose: docs, policies, articles
800–1,200 tokensPreserves long arguments and proceduresTopic dilution; wastes context budget on irrelevant textNarrative documents, legal reasoning, tutorials

Overlap exists to stop answers being guillotined at a boundary. If a chunk ends mid-explanation, neither the chunk before nor the chunk after contains the complete answer. Overlapping by 10–20% of the chunk size means every sentence appears in at least one chunk with its surrounding context.

The cost of overlap is real and worth computing. A 100,000-token corpus split into 500-token chunks with no overlap gives 200 chunks. With 100 tokens of overlap, each chunk advances only 400 tokens, so you get ceil(100000 / 400) = 250 chunks — 25% more vectors to store, embed and search. Push overlap to 250 (50%) and you get 400 chunks, double the index, for diminishing returns.

Overlap is insurance against boundary cuts, and like all insurance it has a premium. Ten to twenty per cent buys most of the protection; fifty per cent mostly buys you a bigger bill.

Splitting should respect structure, not character counts. Break on paragraph breaks first, then sentences, then words — never mid-word. That is exactly what a recursive splitter does.

Decision 2 — Embedding strategy

Two rules here that people break constantly.

The same model must embed both queries and chunks. Different models produce vectors in incompatible spaces. Comparing a vector from model A with a vector from model B gives numbers that look like similarities and mean nothing. If you change embedding model, you must re-embed the entire corpus. There is no migration path.

Many models are asymmetric and require prefixes. A question ("how do I reset my password?") and a passage ("Password resets are handled via the account settings page...") do not look alike lexically or structurally. Models such as the E5 and BGE families are trained with instruction prefixes to bridge this — "query: " for questions, "passage: " for documents. Forgetting the prefix silently drops retrieval quality by several points, with no error message. Check the model card; this is the single most common silent misconfiguration in RAG.

Decision 3 — Top-k

How many chunks to hand the reader. This is a direct recall-versus-noise trade, and it is measurable.

kRecall on a 200-question test set (illustrative run)Context tokens (at 400/chunk)Effect on the reader
10.62400Cheapest; fails outright 38% of the time
30.811,200Good default for narrow, factual questions
50.882,000The usual sweet spot
100.934,000Diminishing returns begin; distractors accumulate
200.958,000Answer quality often drops despite higher recall

That last row is the counter-intuitive one and it catches people out. Going from k=10 to k=20 lifted recall by 2 points but doubled the context and buried the correct passage among nineteen distractors. Models degrade when relevant information sits in the middle of a long context — the "lost in the middle" effect — and every irrelevant chunk is a chance for the model to answer a slightly different question than the one asked.

The way out of this trade is not a bigger k. It is retrieving wide and then re-ranking down: fetch 50, score them properly, keep 5. You get the recall of k=50 with the context size of k=5.

Decision 4 — Similarity metric

Three metrics show up, and the relationship between them matters more than the choice.

Cosine similarity measures the angle between vectors, ignoring their lengths:

cos(a,b)=a⋅b∥a∥ ∥b∥\text{cos}(a,b) = \frac{a \cdot b}{\lVert a \rVert \, \lVert b \rVert}

Dot product is the numerator alone — it rewards both alignment and magnitude. Euclidean distance measures straight-line separation.

The practical point: if all vectors are normalised to unit length, cosine and dot product give identical rankings, and Euclidean distance gives the same ordering too. Concretely, for unit vectors, squared Euclidean distance equals 2 − 2·cos. So a cosine of 0.9 corresponds to a squared distance of 2 − 1.8 = 0.2, and a cosine of 0.5 to 2 − 1.0 = 1.0. Higher cosine, lower distance, same order.

This is why the standard recipe is: normalise your vectors, then use inner product. It is the fastest operation, and it is exactly cosine similarity. The one time it matters is if you do not normalise and use raw dot product — then long vectors win regardless of relevance, and documents with unusual token distributions dominate every result list for no good reason.

Building the retriever

Python
import numpy as np, faissfrom sentence_transformers import SentenceTransformerclass Retriever:    def __init__(self, model_name="BAAI/bge-small-en-v1.5"):        self.model = SentenceTransformer(model_name)        self.chunks = []        self.index = None    def _encode(self, texts, is_query=False):        # BGE is asymmetric: queries get an instruction prefix, passages do not        prefix = "Represent this sentence for searching relevant passages: "        texts = [prefix + t for t in texts] if is_query else texts        v = self.model.encode(texts, normalize_embeddings=True)        return np.asarray(v, dtype="float32")    def build(self, chunks):        self.chunks = chunks        vecs = self._encode([c["text"] for c in chunks])        self.index = faiss.IndexFlatIP(vecs.shape[1])   # inner product == cosine        self.index.add(vecs)    def search(self, query, k=5, min_score=0.0, where=None):        scores, ids = self.index.search(self._encode([query], is_query=True), k * 4)        out = []        for s, i in zip(scores[0], ids[0]):            if i == -1 or s < min_score:                continue            c = self.chunks[i]            if where and not all(c["meta"].get(f) == v for f, v in where.items()):                continue            out.append({**c, "score": float(s)})            if len(out) == k:                break        return out

Two details worth copying. The retriever over-fetches (k * 4) so that score thresholds and metadata filters do not leave you with fewer results than requested. And filtering happens after the vector search here — fine for small corpora, but at scale you want the vector database to apply the filter during search, otherwise a narrow filter over millions of vectors returns nothing at all.

Measuring the retriever on its own

Evaluate the retriever separately from the reader. Otherwise you cannot tell which half is broken. You need a set of questions each labelled with the chunk id that answers it.

MetricQuestion it answersFormula
Recall@k / Hit rateDid the right chunk appear in the top k at all?fraction of queries where a relevant chunk is in the top k
MRRHow near the top was the first right chunk?mean of 1/rank of first relevant hit
Precision@kHow much of what we returned was useful?relevant hits in top k, divided by k

A worked MRR. Five test questions; the rank at which the correct chunk first appeared was 1, 3, 2, not-found, 1. The reciprocal ranks are 1.000, 0.333, 0.500, 0.000, 1.000. Their sum is 2.833, divided by 5 gives MRR = 0.567. Recall@5 for the same run is 4/5 = 0.80.

Read those two numbers together. Recall 0.80 says the retriever usually finds the answer. MRR 0.567 says it often buries it at rank 2 or 3. That combination is the exact signature of a system that needs a re-ranker rather than a better embedding model — the candidates are there, the ordering is wrong.

Designing the reader

The reader takes the question and the surviving chunks and produces an answer. There are two families, and they are genuinely different tools.

Extractive readers: point at a span

An extractive reader selects a contiguous span of the source text as the answer. Classic BERT-style QA models do this by predicting two probability distributions over token positions — where the answer starts and where it ends — and taking the span that maximises the combined score.

Python
import torchfrom transformers import AutoTokenizer, AutoModelForQuestionAnswering# transformers 5 removed the "question-answering" pipeline; call the model directlyname = "deepset/roberta-base-squad2"tok = AutoTokenizer.from_pretrained(name)model = AutoModelForQuestionAnswering.from_pretrained(name)def extract(question, context):    enc = tok(question, context, return_tensors="pt", return_offsets_mapping=True)    offsets = enc.pop("offset_mapping")[0]    with torch.no_grad():        out = model(**enc)    in_ctx = torch.tensor([s == 1 for s in enc.sequence_ids(0)])  # context tokens only    p_start = out.start_logits[0].masked_fill(~in_ctx, -1e9).softmax(-1)    p_end = out.end_logits[0].masked_fill(~in_ctx, -1e9).softmax(-1)    s = int(p_start.argmax())    e = s + int(p_end[s:].argmax())               # the end cannot come before the start    a, b = int(offsets[s][0]), int(offsets[e][1])    return {"answer": context[a:b], "score": round(float(p_start[s] * p_end[e]), 2),            "start": a, "end": b}ctx = ("Enterprise annual contracts carry a 14-day refund window "       "measured from the invoice date. Monthly plans carry 30 days.")print(extract("What is the refund window for enterprise annual contracts?", ctx))# {'answer': '14-day', 'score': 0.87, 'start': 36, 'end': 42}

The span is guaranteed to exist verbatim in the source, so fabrication is impossible by construction. The span score is a real probability from the model, and a low one is a useful signal to abstain — though, like any model score, its threshold has to be tuned on your own data. It runs in about 20 ms on CPU. But it cannot synthesise across two passages, cannot rephrase, cannot say "there are three conditions and they are...", and it fails whenever the answer is not a literal substring.

Generative readers: write the answer

A generative reader is a language model given the context and told to answer from it. It can combine facts from several chunks, resolve pronouns, reformat, summarise, and handle conversational phrasing. It is what almost everyone means by RAG today. The price is that it can fabricate, its confidence is not calibrated, it costs cents rather than nothing, and it takes over a second.

ExtractiveGenerative
OutputExact span from the sourceFreely written text
Can hallucinateNo, structurally impossibleYes, must be constrained by prompt
Multi-passage synthesisNoYes
Confidence scoreA span probability you can thresholdNone you can trust
Latency / cost~20 ms, free after hosting~1,300 ms, per-token pricing
Choose whenAnswers are literal values; audit trail is paramount; huge query volumeAnswers need explanation, synthesis or conversation

Reader design decisions

Keep sampling randomness low. Any randomness in a grounded-answering task is pure downside — you get variability with no creativity benefit, and you lose reproducibility when debugging. Set temperature to 0 where the model accepts it. Some current reasoning models reject a temperature setting or only allow the default; with those, pin the exact model version and let the prompt contract below do the constraining.

Number the chunks and demand citations. Citations are not a nicety; they are the mechanism by which a human can falsify the answer. Numbering also lets you programmatically check that every cited index actually existed.

Order chunks deliberately. Given the lost-in-the-middle effect, one effective trick is to place the highest-scoring chunk first and the second-highest last, pushing weak candidates into the middle where the model attends least.

Make abstention easy and specific. A vague instruction to "say if you don't know" is usually ignored. An exact required string is not.

Python
READER_PROMPT = """You answer strictly from the numbered passages below.Rules:1. Every factual claim ends with its passage number, e.g. [2].2. If the passages do not contain the answer, reply with exactly:   NOT_IN_CONTEXT3. If two passages conflict, report both and cite each.4. Do not add information from your own knowledge.Passages:{context}Question: {question}Answer:"""def read(client, question, chunks):    # strongest first, second-strongest last, weak ones buried in the middle    ordered = chunks[:1] + chunks[2:] + chunks[1:2] if len(chunks) > 2 else chunks    ctx = "\n\n".join(        f"[{i+1}] (source: {c['meta']['source']})\n{c['text']}"        for i, c in enumerate(ordered))    resp = client.chat.completions.create(        model=CHAT_MODEL,                  # from config, as in the first lesson        messages=[{"role": "user",                   "content": READER_PROMPT.format(context=ctx, question=question)}])    text = resp.choices[0].message.content.strip()    if text == "NOT_IN_CONTEXT":        return {"answer": None, "sources": []}    used = sorted({int(n) for n in __import__("re").findall(r"\[(\d+)\]", text)})    return {"answer": text,            "sources": [ordered[n-1]["meta"]["source"] for n in used                        if 1 <= n <= len(ordered)]}

Where this pattern goes wrong

SymptomActual causeFix
Answer is confidently wrong and cites nothingReader ignored the grounding instruction; no abstention pathExact abstention string; reject uncited claims in post-processing
"I don't have that information" on questions you know are coveredRetriever failure — right chunk never surfacedMeasure Recall@k; check embedding prefixes; add a re-ranker
Answers are half-sentences or miss a conditionChunks too small; answer split across a boundaryIncrease chunk size or overlap; use parent-document expansion
Retrieval scores all cluster around 0.7–0.8, right and wrong alikeTopic dilution from oversized chunksShrink chunks; split on structure not character count
Quality collapsed overnight with no code changeEmbedding model version changed; index and queries now in different spacesPin the model version; re-embed the whole corpus on any change
Works in testing, wrong in production on recent topicsIndex is stale relative to the documentsIncremental re-indexing on document change events

The two entries worth dwelling on are the mirror images of each other. A retriever failure looks like a reader failure and vice versa, and teams routinely spend a week tuning prompts when the correct chunk was never in the context at all. The diagnostic takes thirty seconds: log the retrieved chunks for a failing query and read them yourself. If the answer is not in them, stop touching the prompt.

Before you tune the reader, read what the retriever gave it. Most "the model is hallucinating" bug reports are "the model was handed the wrong five paragraphs" bug reports.

Wiring both halves together

Python
class RetrieverReaderQA:    def __init__(self, retriever, client, k=5, min_score=0.30):        self.retriever, self.client = retriever, client        self.k, self.min_score = k, min_score    def ask(self, question, where=None):        hits = self.retriever.search(question, k=self.k,                                     min_score=self.min_score, where=where)        if not hits:            return {"answer": None, "reason": "no_candidates", "sources": []}        # a weak best hit is a strong signal that the corpus does not cover this        if hits[0]["score"] < 0.45:            return {"answer": None, "reason": "low_confidence",                    "top_score": hits[0]["score"], "sources": []}        result = read(self.client, question, hits)        result["retrieval_scores"] = [round(h["score"], 3) for h in hits]        return result

The low_confidence branch is the cheapest quality win available. If the best chunk in a 61,000-chunk index only scores 0.42 against the query, the corpus almost certainly does not cover the question, and calling the reader will cost you money to produce a plausible fabrication. Refusing before generation is faster, cheaper and more honest. Calibrate the threshold on your own data — score distributions differ sharply between embedding models, so a number copied from a blog post will be wrong for you.

Making it better once it works

Two-stage retrieval

Fetch 50 candidates with the fast bi-encoder, score all 50 with a cross-encoder that reads query and passage together, keep the top 5. It costs 50–150 ms and is the direct fix for the "high recall, low MRR" signature in the worked example above: the right chunk is already in the pile, and the cross-encoder moves it to the top. This is the highest-value single change in most systems.

Parent-document retrieval

Index small chunks (200 tokens) for precise matching, but when a chunk is retrieved, give the reader its parent — the 1,000-token section it came from. You match precisely and read with full context, which resolves the classic "chunk was too small to contain the whole condition" failure without inflating the index.

Sentence-window retrieval

The same idea at finer granularity: index individual sentences, but expand each retrieved sentence to include the three sentences either side before handing it to the reader.

Contextual retrieval

A chunk that says "The window is 14 days" has lost the fact that it came from the enterprise annual contract, not the monthly plan. Contextual retrieval, a technique Anthropic published in September 2024, repairs this at indexing time: for each chunk, a model reads the whole document and writes a short note (about 50–100 tokens) situating the chunk — "From the 2025 enterprise terms, section 9, on refunds for annual contracts". That note is prepended to the chunk before it is embedded and before it goes into the keyword index. In Anthropic's own tests, this cut top-20 retrieval failures by 35% with contextual embeddings alone, 49% when combined with contextual BM25, and 67% with re-ranking added on top. The cost is one model call per chunk at indexing time, which prompt caching of the shared document makes much cheaper.

Metadata filtering

Restrict the search space before similarity is computed — department, date range, document status, the requesting user's access level. Filtering out superseded documents is not an optimisation, it is a correctness requirement: a retired policy reads exactly like a current one and will score just as highly.

Caching

Support questions repeat heavily. Cache the embedding of every query (deterministic, so caching is free correctness-wise) and cache full answers for exact repeats with a short expiry.

What this means when you build one

Design the two halves as separate services with separate tests, even if they run in the same process. The retriever's contract is "given a question, return candidate passages with scores"; the reader's contract is "given a question and passages, return a grounded answer or abstain". Keep that boundary clean and every future decision gets easier: you can swap the embedding model without touching the prompt, add a re-ranker as a middle stage, or replace the generative reader with an extractive one for a high-volume endpoint.

Instrument the boundary. Log, for every query, the retrieved chunk ids and scores alongside the final answer. This one log line turns debugging from guesswork into a lookup — when a user reports a bad answer you can see immediately whether the right passage was even present. Teams that skip this end up arguing about prompts for weeks.

Finally, set your thresholds from measurement, not intuition. Build a set of 50 real questions with the correct chunk labelled for each — an afternoon's work. Run the retriever alone and compute Recall@5 and MRR. Recall tells you the ceiling on your whole system's accuracy; no reader can exceed it. MRR tells you whether you need better embeddings (low recall and low MRR) or just better ordering (high recall, low MRR). Those two numbers will direct your next month of work more reliably than any amount of reading individual outputs and forming impressions.