Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
RAG Overview and Pipeline Design
A logistics company put a language model behind their internal help desk. On day three, a warehouse supervisor asked it: "What is the refund window for enterprise annual contracts?"
The model answered, in confident, well-formatted prose: "Enterprise annual contracts carry a 30-day refund window from the invoice date."
The real answer was 14 days. The policy had been tightened in a board meeting eleven weeks earlier. The model had never seen the new policy — it had never seen any of this company's policies. It had seen thousands of SaaS refund pages during training, most of which say 30 days, and it produced the statistically most plausible sentence. Nothing in its output signalled uncertainty. There was no citation to check. The supervisor quoted the number to a customer, and the finance team spent two weeks unpicking it.
This failure is not a bug in that particular model, and a bigger model would not have fixed it. It is structural. A language model's knowledge is frozen inside its weights at training time, and those weights contain no copy of your contracts, your tickets, your wiki, or last Tuesday's incident report. Retrieval-Augmented Generation — RAG — is the engineering answer to that structural gap. This lesson builds the whole pipeline from the failure upward.
Why a language model on its own cannot do this job
It is tempting to treat the above as an accuracy problem to be solved with a better model. It is not. There are four distinct failure modes, and each has a different cause.
Hallucination: fluent output is not grounded output
A language model is trained to predict the next token that a competent writer would produce. It is not trained to predict the true next token, because truth is not a signal available in the loss function — only plausibility is. When the model has genuine knowledge, plausible and true coincide. When it does not, plausibility carries on regardless and produces something that reads exactly like knowledge.
This is why hallucinations are dangerous in a way that ordinary software errors are not. A database that lacks a row returns an empty result. A model that lacks a fact returns a confident sentence.
A language model has no internal signal that distinguishes "I recall this" from "I am inventing something that sounds like this". Both come out in the same voice, at the same fluency, with the same punctuation.
Stale knowledge: the cutoff is a wall, not a fade
Training data has an end date. Everything after it simply does not exist for the model. This is not gradual degradation — a model with a mid-2024 cutoff knows nothing at all about a policy changed in 2025, no matter how important that policy is.
Retraining is not a practical fix. A full pre-training run costs millions of dollars and weeks of compute. Fine-tuning is cheaper but still measured in hours and requires rebuilding a dataset every time a fact changes. Neither is a mechanism you can run when someone edits a wiki page at 4pm.
Domain gaps: your data was never public
Models are trained on broadly public text. Your customer records, your internal runbooks, your unlisted product SKUs, your legal agreements, the Slack thread where the architecture decision was actually made — none of it was in the training set, and you would not want it to be. No amount of scaling reaches data the model was never shown.
No attribution: you cannot check the working
Ask a model where an answer came from and it will produce a plausible-looking citation, sometimes to a document that does not exist. Knowledge in a neural network is distributed across billions of weights; there is no row to point at. For regulated work — medical, legal, financial — an unverifiable answer is not usable at all, regardless of whether it happens to be correct.
What RAG actually is
The insight is almost embarrassingly simple. If the model does not know your facts, put your facts in the question.
Language models are extremely good at reading. Give one a paragraph of text and a question about that paragraph, and it will answer accurately, because now the task is reading comprehension rather than recall. Reading comprehension is something the model does reliably; recall of specific private facts is something it cannot do at all.
So the trick is: before answering, go and fetch the handful of paragraphs from your own documents that are most likely to contain the answer, paste them into the prompt, and instruct the model to answer only from those paragraphs.
RAG converts a recall problem, which language models are bad at, into a reading comprehension problem, which they are extremely good at.
Formally: Retrieval-Augmented Generation is an architecture in which a retrieval system selects relevant documents from an external corpus at query time, and those documents are inserted into the model's prompt as grounding context before generation. Three properties follow directly, and they are the whole business case:
| Property | Why it follows | What it buys you |
|---|---|---|
| Knowledge is updatable in seconds | Facts live in an index, not in weights | Edit a document, re-index it, the system is current. No retraining. |
| Answers are attributable | You know exactly which chunks were in the prompt | Every claim can carry a citation a human can open and verify. |
| Private data stays private | Documents are read at inference, never memorised into weights | Access control can be enforced per query, per user. |
The trade you accept in return: latency goes up (you now do a search before you generate), cost per query goes up (you send far more input tokens), and you inherit a brand new failure surface — if retrieval brings back the wrong paragraphs, the model will faithfully answer the wrong question.
The seven components of a RAG pipeline
Every production RAG system, whatever framework it is written in, decomposes into the same seven parts. Four of them run offline, ahead of time; three run per query.
Offline: building the index
1. Document loading. Get raw content out of wherever it lives — PDFs, Confluence, S3, a Postgres table, Zendesk tickets — and into plain text with metadata attached. Metadata matters more than beginners expect: source URL, author, last-modified date, department, access level. You will want to filter and cite on all of these later, and retrofitting metadata means re-processing the entire corpus.
2. Chunking. Split each document into passages of a few hundred tokens. This exists for two reasons. Embedding models have input limits (commonly 512 tokens), and more importantly, an embedding is a single fixed-length vector — one vector cannot faithfully represent a 40-page document, because averaging forty topics together produces a vector that is near nothing in particular.
3. Embedding. Convert each chunk to a vector of floating point numbers — typically 384, 768, or 1,536 dimensions — using an embedding model. The property that makes this work: chunks with similar meaning land near each other in that space, even when they share no words.
4. Indexing. Store the vectors in a structure that supports fast nearest-neighbour search — a vector database such as FAISS, Qdrant, Pinecone, Weaviate, or the pgvector extension in Postgres. Exact search over a million vectors is too slow for interactive use, so these use approximate algorithms (HNSW being the common one) that trade a small amount of recall for roughly hundred-fold speedups.
Online: answering a query
5. Retrieval. Embed the incoming question with the same model used for the chunks, search the index, return the top-k nearest chunks.
6. Augmentation. Assemble the prompt: a system instruction that constrains the model to the provided context, the retrieved chunks with source labels, and the user's question.
7. Generation. Call the language model, get the answer, attach citations derived from the chunk metadata.
Walking one query end to end
Take the question that started this lesson and trace it through a concrete system: 4,000 internal policy documents, chunked into 61,000 passages of roughly 350 tokens each, embedded at 768 dimensions.
Query: "What is the refund window for enterprise annual contracts?"Step 1 Embed query 768-dim vector produced ~18 msStep 2 Vector search over 61,000 chunks (HNSW) top-5 nearest chunks returned ~11 ms 1. billing-policy-v4.md#refunds cosine 0.87 2. enterprise-terms-2025.pdf#sec-9 cosine 0.84 3. billing-policy-v3.md#refunds cosine 0.81 (superseded!) 4. faq-customer-success.md#returns cosine 0.72 5. sales-playbook.md#objections cosine 0.66Step 3 Build prompt system 180 tokens context 5 chunks x ~350 = 1,750 tokens question 12 tokens -------------------------------- input 1,942 tokensStep 4 Generate ~1,350 ms output 96 tokensTotal wall clock: 18 + 11 + 1,350 = 1,379 msTwo things in that trace deserve attention. First, generation dominates latency by two orders of magnitude — 1,350 ms against 29 ms of retrieval. Optimising your vector search from 11 ms to 6 ms is invisible to users; optimising the number of output tokens is not.
Second, and far more important: chunk 3 is the old version of the policy, and it scored 0.81. It is semantically almost identical to the correct chunk, because a superseded policy reads exactly like a current one. Vector similarity cannot tell you which is authoritative. That is a metadata and filtering problem, and if you do not solve it, your beautifully engineered RAG system will confidently quote a document that was retired in March. This is the single most common way real RAG deployments fail.
Four pipeline patterns, and when each earns its complexity
Pattern 1 — Retrieve and generate
The baseline: embed, search, stuff the top-k into the prompt, generate. One vector lookup, one model call. Latency around 1–2 seconds. This handles the majority of factual lookup questions and is where every project should start. Adding complexity before you have measured this baseline means you cannot tell whether the complexity helped.
Pattern 2 — Retrieve, re-rank, generate
Retrieve a wide net (say top-50) with the fast vector search, then score each candidate against the query with a slower, more accurate cross-encoder model, keep the best 5. The vector search optimises for recall; the re-ranker optimises for precision. Adds roughly 50–150 ms. This is usually the highest return-on-effort upgrade in RAG: the correct passage is often already in the top 50 but sitting at rank 12, and the re-ranker is what moves it into the five the model actually reads.
Pattern 3 — Iterative / multi-hop retrieval
Some questions cannot be answered by any single passage. "Which of our EMEA customers on annual contracts also had a P1 incident last quarter?" requires joining two facts that live in two different documents. Multi-hop retrieval runs retrieval, lets the model decide what it still needs, retrieves again, and repeats until it can answer. Cost: 2–5 times the latency and token spend, plus the risk of the model wandering off. Use it only when your query log shows genuine composition.
Pattern 4 — Hierarchical retrieval
Index at two levels: document-level summaries and chunk-level passages. First search the summaries to identify the 3 relevant documents out of 4,000, then search chunks only within those. This keeps precision high on very large or very heterogeneous corpora, where a chunk-level search across everything drowns in near-duplicates.
| Pattern | Model calls | Typical latency | Use when | Do not use when |
|---|---|---|---|---|
| Retrieve → generate | 1 | 1–2 s | Direct factual lookup; you are starting out | Answers need facts from several documents |
| Retrieve → re-rank → generate | 1 + 1 re-ranker | 1.2–2.5 s | Almost always; precision matters | Hard sub-300 ms budget |
| Iterative / multi-hop | 2–5 | 4–12 s | Comparative or compositional questions | Latency-sensitive; simple lookups |
| Hierarchical | 1–2 | 1.5–3 s | Corpora above ~100k chunks; many near-duplicates | Small, homogeneous corpora |
The five design decisions that determine whether it works
Retrieval quality is the ceiling
If the correct passage is not in the retrieved set, no prompt engineering and no larger model will recover the answer. The model cannot read what it was not given. This makes retrieval recall the hard ceiling on end-to-end accuracy, and it is where your measurement effort belongs. When a RAG system answers badly, check retrieval first — in practice a large share of RAG failures are retrieval failures wearing a generation costume.
Context budget is a real budget
Large context windows tempt people into retrieving 50 chunks "just in case". Two things go wrong. Models exhibit a well-documented lost in the middle effect: information placed in the centre of a long context is used measurably less reliably than information at the start or end. And every irrelevant chunk is a distractor that can pull the answer off course.
The cost arithmetic is stark. Suppose 8 chunks of 400 tokens plus 200 tokens of instruction — that is 3,400 input tokens per query. At a rate of 3 dollars per million input tokens, one query costs 0.0102 dollars. At 40,000 queries a day, that is 408 dollars a day, or roughly 12,240 dollars a month, for input alone. Retrieving 24 chunks instead of 8 triples it to about 36,700 dollars a month — and often makes answers worse.
The same arithmetic answers a question teams now ask often: current models accept hundreds of thousands of tokens, so why not skip retrieval and paste in the whole corpus? For a small, stable set of documents — one contract, one manual, a few dozen pages — that can be the simpler and better design, especially with prompt caching. It stops working as the corpus grows. 4,000 policy documents will not fit in any context window, every query would pay for the full corpus in input tokens and latency, and the lost-in-the-middle effect gets worse, not better, as the context grows. You would also lose per-user access control and a clean citation trail. Long context changes where the line sits, not the need for retrieval.
Latency is spent unevenly
From the trace above: embedding 18 ms, search 11 ms, generation 1,350 ms. If you need to get faster, the levers that matter are streaming the response so time-to-first-token drops, caching repeated queries, and shortening the requested output — not micro-tuning the index.
Hallucination control is a prompt contract
Retrieval alone does not stop a model inventing things. You must explicitly instruct it to abstain, and you must make abstaining an acceptable outcome:
1SYSTEM = """Answer the question using ONLY the numbered context passages below.23Rules:4- Cite the passage number for every factual claim, like [2].5- If the passages do not contain the answer, reply exactly:6 "I don't have that information in the available documents."7- Never use knowledge from outside the passages.8- If passages disagree, say so and cite both.910Context:11{numbered_context}1213Question: {question}"""That last rule matters because of the superseded-policy problem. Two chunks saying "30 days" and "14 days" should produce a flagged conflict, not a coin flip.
Scalability changes the architecture, not just the hardware
| Corpus size | Reasonable index | What changes |
|---|---|---|
| Under ~50k chunks | FAISS flat, or pgvector in your existing Postgres | Exact search is fast enough; keep it simple |
| 50k – 5M chunks | HNSW index (Qdrant, Weaviate, pgvector HNSW) | Approximate search; tune ef_search for recall |
| Above ~5M chunks | Sharded / managed service, quantised vectors | Memory becomes the constraint; consider int8 quantisation |
A useful memory estimate: 1 million chunks at 768 dimensions in float32 is 1,000,000 × 768 × 4 bytes ≈ 3.07 GB of raw vectors, before index overhead. HNSW typically adds 30–60% on top. Quantising to int8 cuts the raw figure to about 0.77 GB at a small recall cost.
Building it: frameworks versus writing it yourself
Here is the same pipeline three ways, so you can see what the frameworks are actually doing.
1# --- LangChain 1.x: composable, explicit about each step ---2import os3from langchain_core.vectorstores import InMemoryVectorStore4from langchain_openai import OpenAIEmbeddings, ChatOpenAI5from langchain_text_splitters import RecursiveCharacterTextSplitter67CHAT_MODEL = os.getenv("CHAT_MODEL", "gpt-6-luna") # model names live in config89splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)10chunks = splitter.split_documents(docs)1112# name the embedding model: OpenAIEmbeddings() still defaults to ada-00213emb = OpenAIEmbeddings(model="text-embedding-3-small")14store = InMemoryVectorStore.from_documents(chunks, emb)15retriever = store.as_retriever(search_kwargs={"k": 5})1617def answer(question):18 hits = retriever.invoke(question)19 ctx = "\n\n".join(f"[{i+1}] {d.page_content}" for i, d in enumerate(hits))20 prompt = SYSTEM.format(numbered_context=ctx, question=question)21 return ChatOpenAI(model=CHAT_MODEL).invoke(prompt).content1# --- LlamaIndex: fewer lines, ingestion handled for you ---2from llama_index.core import VectorStoreIndex, SimpleDirectoryReader34docs = SimpleDirectoryReader("./policies").load_data()5index = VectorStoreIndex.from_documents(docs)6engine = index.as_query_engine(similarity_top_k=5)7print(engine.query("Refund window for enterprise annual contracts?"))1# --- No framework: ~30 lines, zero abstraction to debug through ---2import os3import numpy as np, faiss4from openai import OpenAI56client = OpenAI()7CHAT_MODEL = os.getenv("CHAT_MODEL", "gpt-6-luna")89def embed(texts):10 r = client.embeddings.create(model="text-embedding-3-small", input=texts)11 v = np.array([d.embedding for d in r.data], dtype="float32")12 faiss.normalize_L2(v) # cosine similarity via inner product13 return v1415vecs = embed(chunk_texts)16index = faiss.IndexFlatIP(vecs.shape[1])17index.add(vecs)1819def answer(question, k=5):20 scores, ids = index.search(embed([question]), k)21 ctx = "\n\n".join(f"[{i+1}] {chunk_texts[j]}" for i, j in enumerate(ids[0]))22 msg = SYSTEM.format(numbered_context=ctx, question=question)23 out = client.chat.completions.create(24 model=CHAT_MODEL,25 messages=[{"role": "user", "content": msg}])26 return out.choices[0].message.content, scores[0]| LangChain | LlamaIndex | Hand-rolled | |
|---|---|---|---|
| Lines to a prototype | ~20 | ~6 | ~30 |
| Loaders and integrations | Very many | Many, ingestion-focused | You write them |
| Ease of debugging | Moderate — abstraction layers | Moderate | Highest — it is all your code |
| Control over prompt and ranking | Good | Good with effort | Total |
| Best for | Multi-step chains and agents | Fast document-QA builds | Production systems with unusual requirements |
Two notes on the LangChain version. In LangChain 1.x the old all-in-one retrieval chains (RetrievalQA and friends) moved to the langchain-classic package, and langchain-community, where most vector store wrappers used to live, is being sunset in favour of standalone integration packages such as langchain-chroma or langchain-postgres. InMemoryVectorStore ships with langchain-core and is fine for a prototype; swap in a real store's package for production. Retrievers are called with .invoke(question); the older get_relevant_documents method is gone.
A pragmatic route many teams take: prototype in a framework to validate the idea in an afternoon, then rewrite the hot path by hand once you know exactly which knobs you need. The hand-rolled version above is not much code, and when retrieval misbehaves at 2am you can read every line of it.
What this means when you build one
The mistake that wastes the most time on a first RAG project is treating it as a prompt engineering exercise. It is not. It is a search engineering exercise with a language model bolted on the end, and the search half is where almost all the difficulty lives.
Concretely, before you write the pipeline, write down twenty real questions your users will ask, with the passage that correctly answers each one. That is your evaluation set, and it takes about an hour to build. Then measure one number: for what fraction of those twenty questions does the correct passage appear in the retrieved top-5? If it is 0.6, your system's accuracy ceiling is 60% and no prompt will lift it. Fix retrieval — better chunking, a re-ranker, metadata filters — and watch that number, not the vibes of individual answers.
Second, build the metadata schema before you index anything. Source, URL, last-modified date, version, department, access level. The superseded-policy failure in the trace above is fixed by one filter on a status field — but only if that field exists at index time. Re-embedding 61,000 chunks because you forgot to store a date is a genuinely miserable afternoon.
Third, decide now what "I don't know" looks like in your product, and make it a first-class outcome. A system that abstains on 8% of questions and is right on the other 92% is far more valuable than one that answers everything and is right on 85%, because users can trust the first one and cannot trust the second. Grounding is not just about retrieving the right text; it is about being willing to return nothing when the right text is not there.