Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

RAG Overview and Pipeline Design


A logistics company put a language model behind their internal help desk. On day three, a warehouse supervisor asked it: "What is the refund window for enterprise annual contracts?"

The model answered, in confident, well-formatted prose: "Enterprise annual contracts carry a 30-day refund window from the invoice date."

The real answer was 14 days. The policy had been tightened in a board meeting eleven weeks earlier. The model had never seen the new policy — it had never seen any of this company's policies. It had seen thousands of SaaS refund pages during training, most of which say 30 days, and it produced the statistically most plausible sentence. Nothing in its output signalled uncertainty. There was no citation to check. The supervisor quoted the number to a customer, and the finance team spent two weeks unpicking it.

This failure is not a bug in that particular model, and a bigger model would not have fixed it. It is structural. A language model's knowledge is frozen inside its weights at training time, and those weights contain no copy of your contracts, your tickets, your wiki, or last Tuesday's incident report. Retrieval-Augmented Generation — RAG — is the engineering answer to that structural gap. This lesson builds the whole pipeline from the failure upward.

Two pipelines that share only the indexOffline — build the index• Load documents from every source• Chunk with an overlap you chose• Embed each chunk once, in batches• Write vectors plus metadata to the storeOnline — answer one query• Embed the question with the same model• Retrieve the top k nearest chunks• Re-rank, then trim to the context budget• Generate, citing thechunks that were used
The offline path runs once per document and the online path runs once per keystroke, which is why almost all the compute you can afford belongs on the left.

Why a language model on its own cannot do this job

It is tempting to treat the above as an accuracy problem to be solved with a better model. It is not. There are four distinct failure modes, and each has a different cause.

Hallucination: fluent output is not grounded output

A language model is trained to predict the next token that a competent writer would produce. It is not trained to predict the true next token, because truth is not a signal available in the loss function — only plausibility is. When the model has genuine knowledge, plausible and true coincide. When it does not, plausibility carries on regardless and produces something that reads exactly like knowledge.

This is why hallucinations are dangerous in a way that ordinary software errors are not. A database that lacks a row returns an empty result. A model that lacks a fact returns a confident sentence.

A language model has no internal signal that distinguishes "I recall this" from "I am inventing something that sounds like this". Both come out in the same voice, at the same fluency, with the same punctuation.

Stale knowledge: the cutoff is a wall, not a fade

Training data has an end date. Everything after it simply does not exist for the model. This is not gradual degradation — a model with a mid-2024 cutoff knows nothing at all about a policy changed in 2025, no matter how important that policy is.

Retraining is not a practical fix. A full pre-training run costs millions of dollars and weeks of compute. Fine-tuning is cheaper but still measured in hours and requires rebuilding a dataset every time a fact changes. Neither is a mechanism you can run when someone edits a wiki page at 4pm.

Domain gaps: your data was never public

Models are trained on broadly public text. Your customer records, your internal runbooks, your unlisted product SKUs, your legal agreements, the Slack thread where the architecture decision was actually made — none of it was in the training set, and you would not want it to be. No amount of scaling reaches data the model was never shown.

No attribution: you cannot check the working

Ask a model where an answer came from and it will produce a plausible-looking citation, sometimes to a document that does not exist. Knowledge in a neural network is distributed across billions of weights; there is no row to point at. For regulated work — medical, legal, financial — an unverifiable answer is not usable at all, regardless of whether it happens to be correct.

What RAG actually is

The insight is almost embarrassingly simple. If the model does not know your facts, put your facts in the question.

Language models are extremely good at reading. Give one a paragraph of text and a question about that paragraph, and it will answer accurately, because now the task is reading comprehension rather than recall. Reading comprehension is something the model does reliably; recall of specific private facts is something it cannot do at all.

So the trick is: before answering, go and fetch the handful of paragraphs from your own documents that are most likely to contain the answer, paste them into the prompt, and instruct the model to answer only from those paragraphs.

RAG converts a recall problem, which language models are bad at, into a reading comprehension problem, which they are extremely good at.

Formally: Retrieval-Augmented Generation is an architecture in which a retrieval system selects relevant documents from an external corpus at query time, and those documents are inserted into the model's prompt as grounding context before generation. Three properties follow directly, and they are the whole business case:

PropertyWhy it followsWhat it buys you
Knowledge is updatable in secondsFacts live in an index, not in weightsEdit a document, re-index it, the system is current. No retraining.
Answers are attributableYou know exactly which chunks were in the promptEvery claim can carry a citation a human can open and verify.
Private data stays privateDocuments are read at inference, never memorised into weightsAccess control can be enforced per query, per user.

The trade you accept in return: latency goes up (you now do a search before you generate), cost per query goes up (you send far more input tokens), and you inherit a brand new failure surface — if retrieval brings back the wrong paragraphs, the model will faithfully answer the wrong question.

The seven components of a RAG pipeline

Every production RAG system, whatever framework it is written in, decomposes into the same seven parts. Four of them run offline, ahead of time; three run per query.

Offline: building the index

1. Document loading. Get raw content out of wherever it lives — PDFs, Confluence, S3, a Postgres table, Zendesk tickets — and into plain text with metadata attached. Metadata matters more than beginners expect: source URL, author, last-modified date, department, access level. You will want to filter and cite on all of these later, and retrofitting metadata means re-processing the entire corpus.

2. Chunking. Split each document into passages of a few hundred tokens. This exists for two reasons. Embedding models have input limits (commonly 512 tokens), and more importantly, an embedding is a single fixed-length vector — one vector cannot faithfully represent a 40-page document, because averaging forty topics together produces a vector that is near nothing in particular.

3. Embedding. Convert each chunk to a vector of floating point numbers — typically 384, 768, or 1,536 dimensions — using an embedding model. The property that makes this work: chunks with similar meaning land near each other in that space, even when they share no words.

4. Indexing. Store the vectors in a structure that supports fast nearest-neighbour search — a vector database such as FAISS, Qdrant, Pinecone, Weaviate, or the pgvector extension in Postgres. Exact search over a million vectors is too slow for interactive use, so these use approximate algorithms (HNSW being the common one) that trade a small amount of recall for roughly hundred-fold speedups.

Online: answering a query

5. Retrieval. Embed the incoming question with the same model used for the chunks, search the index, return the top-k nearest chunks.

6. Augmentation. Assemble the prompt: a system instruction that constrains the model to the provided context, the retrieved chunks with source labels, and the user's question.

7. Generation. Call the language model, get the answer, attach citations derived from the chunk metadata.

Walking one query end to end

Take the question that started this lesson and trace it through a concrete system: 4,000 internal policy documents, chunked into 61,000 passages of roughly 350 tokens each, embedded at 768 dimensions.

Text
Query: "What is the refund window for enterprise annual contracts?"Step 1  Embed query        768-dim vector produced          ~18 msStep 2  Vector search over 61,000 chunks (HNSW)        top-5 nearest chunks returned    ~11 ms        1. billing-policy-v4.md#refunds       cosine 0.87        2. enterprise-terms-2025.pdf#sec-9    cosine 0.84        3. billing-policy-v3.md#refunds       cosine 0.81   (superseded!)        4. faq-customer-success.md#returns    cosine 0.72        5. sales-playbook.md#objections       cosine 0.66Step 3  Build prompt        system   180 tokens        context  5 chunks x ~350 = 1,750 tokens        question  12 tokens        --------------------------------        input    1,942 tokensStep 4  Generate                          ~1,350 ms        output    96 tokensTotal wall clock: 18 + 11 + 1,350 = 1,379 ms

Two things in that trace deserve attention. First, generation dominates latency by two orders of magnitude — 1,350 ms against 29 ms of retrieval. Optimising your vector search from 11 ms to 6 ms is invisible to users; optimising the number of output tokens is not.

Second, and far more important: chunk 3 is the old version of the policy, and it scored 0.81. It is semantically almost identical to the correct chunk, because a superseded policy reads exactly like a current one. Vector similarity cannot tell you which is authoritative. That is a metadata and filtering problem, and if you do not solve it, your beautifully engineered RAG system will confidently quote a document that was retired in March. This is the single most common way real RAG deployments fail.

Four pipeline patterns, and when each earns its complexity

Pattern 1 — Retrieve and generate

The baseline: embed, search, stuff the top-k into the prompt, generate. One vector lookup, one model call. Latency around 1–2 seconds. This handles the majority of factual lookup questions and is where every project should start. Adding complexity before you have measured this baseline means you cannot tell whether the complexity helped.

Pattern 2 — Retrieve, re-rank, generate

Retrieve a wide net (say top-50) with the fast vector search, then score each candidate against the query with a slower, more accurate cross-encoder model, keep the best 5. The vector search optimises for recall; the re-ranker optimises for precision. Adds roughly 50–150 ms. This is usually the highest return-on-effort upgrade in RAG: the correct passage is often already in the top 50 but sitting at rank 12, and the re-ranker is what moves it into the five the model actually reads.

Pattern 3 — Iterative / multi-hop retrieval

Some questions cannot be answered by any single passage. "Which of our EMEA customers on annual contracts also had a P1 incident last quarter?" requires joining two facts that live in two different documents. Multi-hop retrieval runs retrieval, lets the model decide what it still needs, retrieves again, and repeats until it can answer. Cost: 2–5 times the latency and token spend, plus the risk of the model wandering off. Use it only when your query log shows genuine composition.

Pattern 4 — Hierarchical retrieval

Index at two levels: document-level summaries and chunk-level passages. First search the summaries to identify the 3 relevant documents out of 4,000, then search chunks only within those. This keeps precision high on very large or very heterogeneous corpora, where a chunk-level search across everything drowns in near-duplicates.

PatternModel callsTypical latencyUse whenDo not use when
Retrieve → generate11–2 sDirect factual lookup; you are starting outAnswers need facts from several documents
Retrieve → re-rank → generate1 + 1 re-ranker1.2–2.5 sAlmost always; precision mattersHard sub-300 ms budget
Iterative / multi-hop2–54–12 sComparative or compositional questionsLatency-sensitive; simple lookups
Hierarchical1–21.5–3 sCorpora above ~100k chunks; many near-duplicatesSmall, homogeneous corpora

The five design decisions that determine whether it works

Retrieval quality is the ceiling

If the correct passage is not in the retrieved set, no prompt engineering and no larger model will recover the answer. The model cannot read what it was not given. This makes retrieval recall the hard ceiling on end-to-end accuracy, and it is where your measurement effort belongs. When a RAG system answers badly, check retrieval first — in practice a large share of RAG failures are retrieval failures wearing a generation costume.

Context budget is a real budget

Large context windows tempt people into retrieving 50 chunks "just in case". Two things go wrong. Models exhibit a well-documented lost in the middle effect: information placed in the centre of a long context is used measurably less reliably than information at the start or end. And every irrelevant chunk is a distractor that can pull the answer off course.

The cost arithmetic is stark. Suppose 8 chunks of 400 tokens plus 200 tokens of instruction — that is 3,400 input tokens per query. At a rate of 3 dollars per million input tokens, one query costs 0.0102 dollars. At 40,000 queries a day, that is 408 dollars a day, or roughly 12,240 dollars a month, for input alone. Retrieving 24 chunks instead of 8 triples it to about 36,700 dollars a month — and often makes answers worse.

The same arithmetic answers a question teams now ask often: current models accept hundreds of thousands of tokens, so why not skip retrieval and paste in the whole corpus? For a small, stable set of documents — one contract, one manual, a few dozen pages — that can be the simpler and better design, especially with prompt caching. It stops working as the corpus grows. 4,000 policy documents will not fit in any context window, every query would pay for the full corpus in input tokens and latency, and the lost-in-the-middle effect gets worse, not better, as the context grows. You would also lose per-user access control and a clean citation trail. Long context changes where the line sits, not the need for retrieval.

Latency is spent unevenly

From the trace above: embedding 18 ms, search 11 ms, generation 1,350 ms. If you need to get faster, the levers that matter are streaming the response so time-to-first-token drops, caching repeated queries, and shortening the requested output — not micro-tuning the index.

Hallucination control is a prompt contract

Retrieval alone does not stop a model inventing things. You must explicitly instruct it to abstain, and you must make abstaining an acceptable outcome:

Python
SYSTEM = """Answer the question using ONLY the numbered context passages below.Rules:- Cite the passage number for every factual claim, like [2].- If the passages do not contain the answer, reply exactly:  "I don't have that information in the available documents."- Never use knowledge from outside the passages.- If passages disagree, say so and cite both.Context:{numbered_context}Question: {question}"""

That last rule matters because of the superseded-policy problem. Two chunks saying "30 days" and "14 days" should produce a flagged conflict, not a coin flip.

Scalability changes the architecture, not just the hardware

Corpus sizeReasonable indexWhat changes
Under ~50k chunksFAISS flat, or pgvector in your existing PostgresExact search is fast enough; keep it simple
50k – 5M chunksHNSW index (Qdrant, Weaviate, pgvector HNSW)Approximate search; tune ef_search for recall
Above ~5M chunksSharded / managed service, quantised vectorsMemory becomes the constraint; consider int8 quantisation

A useful memory estimate: 1 million chunks at 768 dimensions in float32 is 1,000,000 × 768 × 4 bytes ≈ 3.07 GB of raw vectors, before index overhead. HNSW typically adds 30–60% on top. Quantising to int8 cuts the raw figure to about 0.77 GB at a small recall cost.

Building it: frameworks versus writing it yourself

Here is the same pipeline three ways, so you can see what the frameworks are actually doing.

Python
# --- LangChain 1.x: composable, explicit about each step ---import osfrom langchain_core.vectorstores import InMemoryVectorStorefrom langchain_openai import OpenAIEmbeddings, ChatOpenAIfrom langchain_text_splitters import RecursiveCharacterTextSplitterCHAT_MODEL = os.getenv("CHAT_MODEL", "gpt-6-luna")   # model names live in configsplitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=150)chunks = splitter.split_documents(docs)# name the embedding model: OpenAIEmbeddings() still defaults to ada-002emb = OpenAIEmbeddings(model="text-embedding-3-small")store = InMemoryVectorStore.from_documents(chunks, emb)retriever = store.as_retriever(search_kwargs={"k": 5})def answer(question):    hits = retriever.invoke(question)    ctx = "\n\n".join(f"[{i+1}] {d.page_content}" for i, d in enumerate(hits))    prompt = SYSTEM.format(numbered_context=ctx, question=question)    return ChatOpenAI(model=CHAT_MODEL).invoke(prompt).content
Python
# --- LlamaIndex: fewer lines, ingestion handled for you ---from llama_index.core import VectorStoreIndex, SimpleDirectoryReaderdocs = SimpleDirectoryReader("./policies").load_data()index = VectorStoreIndex.from_documents(docs)engine = index.as_query_engine(similarity_top_k=5)print(engine.query("Refund window for enterprise annual contracts?"))
Python
# --- No framework: ~30 lines, zero abstraction to debug through ---import osimport numpy as np, faissfrom openai import OpenAIclient = OpenAI()CHAT_MODEL = os.getenv("CHAT_MODEL", "gpt-6-luna")def embed(texts):    r = client.embeddings.create(model="text-embedding-3-small", input=texts)    v = np.array([d.embedding for d in r.data], dtype="float32")    faiss.normalize_L2(v)          # cosine similarity via inner product    return vvecs = embed(chunk_texts)index = faiss.IndexFlatIP(vecs.shape[1])index.add(vecs)def answer(question, k=5):    scores, ids = index.search(embed([question]), k)    ctx = "\n\n".join(f"[{i+1}] {chunk_texts[j]}" for i, j in enumerate(ids[0]))    msg = SYSTEM.format(numbered_context=ctx, question=question)    out = client.chat.completions.create(        model=CHAT_MODEL,        messages=[{"role": "user", "content": msg}])    return out.choices[0].message.content, scores[0]
LangChainLlamaIndexHand-rolled
Lines to a prototype~20~6~30
Loaders and integrationsVery manyMany, ingestion-focusedYou write them
Ease of debuggingModerate — abstraction layersModerateHighest — it is all your code
Control over prompt and rankingGoodGood with effortTotal
Best forMulti-step chains and agentsFast document-QA buildsProduction systems with unusual requirements

Two notes on the LangChain version. In LangChain 1.x the old all-in-one retrieval chains (RetrievalQA and friends) moved to the langchain-classic package, and langchain-community, where most vector store wrappers used to live, is being sunset in favour of standalone integration packages such as langchain-chroma or langchain-postgres. InMemoryVectorStore ships with langchain-core and is fine for a prototype; swap in a real store's package for production. Retrievers are called with .invoke(question); the older get_relevant_documents method is gone.

A pragmatic route many teams take: prototype in a framework to validate the idea in an afternoon, then rewrite the hot path by hand once you know exactly which knobs you need. The hand-rolled version above is not much code, and when retrieval misbehaves at 2am you can read every line of it.

What this means when you build one

The mistake that wastes the most time on a first RAG project is treating it as a prompt engineering exercise. It is not. It is a search engineering exercise with a language model bolted on the end, and the search half is where almost all the difficulty lives.

Concretely, before you write the pipeline, write down twenty real questions your users will ask, with the passage that correctly answers each one. That is your evaluation set, and it takes about an hour to build. Then measure one number: for what fraction of those twenty questions does the correct passage appear in the retrieved top-5? If it is 0.6, your system's accuracy ceiling is 60% and no prompt will lift it. Fix retrieval — better chunking, a re-ranker, metadata filters — and watch that number, not the vibes of individual answers.

Second, build the metadata schema before you index anything. Source, URL, last-modified date, version, department, access level. The superseded-policy failure in the trace above is fixed by one filter on a status field — but only if that field exists at index time. Re-embedding 61,000 chunks because you forgot to store a date is a genuinely miserable afternoon.

Third, decide now what "I don't know" looks like in your product, and make it a first-class outcome. A system that abstains on 8% of questions and is right on the other 92% is far more valuable than one that answers everything and is right on 85%, because users can trust the first one and cannot trust the second. Grounding is not just about retrieving the right text; it is about being willing to return nothing when the right text is not there.