Retrieval-Augmented Generation (RAG)

Course Content

Retrieval-Augmented Generation (RAG)

4 sections · 8 lessons

Mini Project: Domain-Specific RAG Chatbot


Almost everyone's first RAG chatbot demos beautifully and then falls over in public.

The pattern is consistent. You load 40 PDFs, split them at 1,000 characters, embed them, wire up a retriever and a prompt, and ask five questions you already know the answers to. All five come back right. You show someone. They ask a sixth question — "does that apply to contractors as well?" — and the bot invents a policy that does not exist, in the same confident tone it used for the five correct answers.

Nothing malfunctioned. The retriever returned its top 5 chunks, none of which mentioned contractors, and the model filled the gap because nothing in the prompt told it not to. The five questions you tested with were the five you had unconsciously written to match your chunking.

This project is the fix. You will build a domain-specific RAG chatbot end to end, and — the part that matters — you will measure it honestly enough to know which stage is failing and what each change actually bought you. Budget about four and a half hours.

Chunking with overlap, in tokens0641281922563203844485125760123456789chunk 1startschunk 2startschunk 1 ends512-token chunks with a 64-token stride of overlap; the boxed region belongs to both.
Overlap exists so a sentence that straddles a boundary is whole in at least one chunk — the cost is storing and embedding roughly 12 percent more text.

What you are building

A question-answering system over a document collection you choose, meeting concrete acceptance criteria rather than "it works":

RequirementTarget
Corpus size30-100 documents, 150,000+ words total
Retrieval qualityRecall@5 ≥ 0.85 on your test set
Faithfulness≥ 0.90 — claims supported by retrieved context
Refusal behaviourCorrectly declines ≥ 80% of unanswerable questions
CitationsEvery answer names its source documents
Latencyp95 under 5 seconds end to end
EvaluationScripted, reproducible, reports retrieval and generation separately

The refusal row is what separates a real system from a demo. A chatbot that answers everything scores perfectly on a test set where everything is answerable, and hallucinates the moment it meets a real user.

Data preparation — 45 minutes

Choose a domain the model does not already know

This decision determines whether the whole project teaches you anything, and most people get it wrong.

Pick "Python documentation" or "the history of the Roman Empire" and your chatbot will answer well with the retriever completely disconnected, because the model already knows the material. You will have built nothing and measured nothing.

Run the control test before you write a line of pipeline code. Take 20 questions from your intended domain and ask a bare language model with no retrieval at all.

Bare model gets rightVerdict
15-20 of 20Wrong domain. The model knows it. Choose again.
6-14 of 20Usable, but your measured gains will be muddied.
0-5 of 20Correct choice. Every point of accuracy came from your system.

Domains that reliably pass: your own organisation's internal documentation; a niche regulatory corpus; product manuals for specific equipment; a set of research papers from the last twelve months; the complete rules of a game with a small following; the minutes and bylaws of a local institution.

Collect and validate

Aim for 30-100 documents totalling 150,000 words or more — roughly 200 pages. Below that, retrieval is trivially easy and you will not see the failure modes the project exists to teach.

Validate before indexing, because every extraction problem becomes a retrieval problem you will misdiagnose for an hour:

Python
from pathlib import Pathdef validate(docs):    report = []    for d in docs:        text = d.page_content        words = len(text.split())        alpha = sum(c.isalpha() for c in text) / max(len(text), 1)        report.append({            "source":   d.metadata["source"],            "words":    words,            "alpha":    round(alpha, 2),   # < 0.6 suggests bad extraction            "empty":    words < 50,            "no_space": " " not in text[:200],  # PDF ligature failure        })    return report

The two checks that catch the most damage are the alphabetic ratio and the missing-space test. A PDF that extracted as "Refundswithinfourteendays" looks fine in a file listing, embeds to something meaningless, and silently removes that document from your corpus. You will spend an hour blaming the embedding model.

Attach metadata now, not later: source, doc_id, title, section, date, and any category you might want to filter on. Retrofitting metadata means re-embedding everything.

Retriever implementation — 60 minutes

Chunking is the decision that matters most

Most beginner RAG failures are chunking failures wearing a costume. A chunk that splits a policy from its exception produces a retriever that confidently returns half a rule.

Chunk sizeChunks from 200k words (15% overlap)Recall behaviourFailure mode
200 tokens~1,600Sharp, focused embeddingsAnswers split across chunks; no chunk is complete
400 tokens~800Best general-purpose choiceOccasional split at section boundaries
800 tokens~400Whole sections survive intactDiluted embeddings; one chunk covers three topics
1,500 tokens~215Nothing gets splitRetrieval becomes near-random; huge context cost

Start at 400 tokens with 15% overlap. The overlap exists so that a sentence straddling a boundary appears whole in at least one chunk — without it, "…the window is 14 days. | Contractors are excluded…" produces two chunks, neither of which answers "how long do contractors get?"

Python
from langchain_text_splitters import RecursiveCharacterTextSplittersplitter = RecursiveCharacterTextSplitter(    chunk_size=1600,            # ~400 tokens    chunk_overlap=240,          # 15%    separators=["\n## ", "\n### ", "\n\n", "\n", ". ", " "],)chunks = splitter.split_documents(docs)

The separators list is doing quiet, important work. It tries to break at a heading first, then a paragraph, then a line, then a sentence, and only mid-word as a last resort. Left at the default, you get splits in the middle of sentences and a measurable recall drop.

Index and configure

Embedding 200,000 words is around 270,000 tokens — a few pence and under two minutes. Cost is not a constraint at this scale, so use a good embedding model rather than a cheap one.

Configure the retriever to fetch wide, not narrow. Fetch 30 candidates and send 5 to the generator. Retrieving 5 directly is the most common configuration mistake in beginner projects: it gives the re-ranker nothing to work with and puts a hard ceiling on accuracy equal to your recall@5.

Python
retriever = store.as_retriever(search_kwargs={"k": 30})def get_context(query, k=5):    candidates = retriever.invoke(query)    scored = reranker.predict([(query, c.page_content) for c in candidates])    ranked = sorted(zip(scored, candidates), key=lambda x: x[0], reverse=True)    return [d for _, d in ranked[:k]]

Add BM25 keyword search alongside the vector search and merge the two result lists. Every real corpus contains error codes, part numbers, statute references and proper nouns that embeddings handle badly, because those tokens carry no learned meaning. Keyword search finds them exactly.

Generator integration — 60 minutes

The prompt is where hallucination is prevented

Compare two prompts that look equally reasonable.

Text
WEAK:Answer the question using the context below.Context: {context}Question: {question}
Text
STRONG:You answer questions using ONLY the provided context.Rules:1. Every factual statement must be supported by the context. Do not infer,   generalise, or fill gaps from your own knowledge.2. Cite the source after each claim, as [source: filename].3. If the context does not contain the answer, reply exactly:   "I could not find this in the available documents." Then name the   closest topics the documents do cover.4. If sources disagree, say so and quote both.Context:{context}Question: {question}Answer:

The weak prompt permits everything the strong one forbids. "Using the context" does not say only the context, does not forbid inference, and offers no escape hatch — so when the context is silent, the model's only available behaviour is to guess. In the ablation below, moving from the weak prompt to the strong one with temperature at 0 was worth twelve points of end-to-end accuracy and cost nothing.

Rule 3 matters more than it looks. Give a model an explicit sentence to say when it does not know, and it will use it. Leave the refusal undefined and it will not invent one.

Wire the chain and put a face on it

Python
def format_context(docs):    return "\n\n".join(        f"[source: {d.metadata['source']}]\n{d.page_content}" for d in docs    )def answer(question, history=None):    standalone = condense(history, question) if history else question    docs = get_context(standalone, k=5)    prompt = STRONG.format(context=format_context(docs), question=standalone)    text = llm.invoke(prompt).content   # llm built once from config; temperature 0 if the model allows it    return text, [d.metadata["source"] for d in docs]

Two details carry weight. Putting the source marker inside the context block is what makes citation possible — the model cannot cite what it cannot see. And condense rewrites follow-up questions into standalone ones, so "does that apply to contractors?" reaches the retriever as "does the 14-day refund window apply to contractors?". Skip that step and every multi-turn conversation retrieves noise from turn two onwards.

The interface can be a command-line loop or a small Streamlit app. What it must show, either way, is the retrieved sources and their scores next to every answer. You cannot debug what you cannot see, and a bot that displays its evidence is the single most useful debugging tool in the project.

Evaluation and optimisation — 75 minutes

This phase carries the largest share of the marks, because it is the part that turns a demo into engineering.

Build the test set first, and stratify it

Thirty questions minimum. For each, record the question, the gold answer, and the document IDs that contain it — that last field is what lets you score retrieval separately from generation.

TypeCountWhat it exposes
Single-fact lookup12Baseline competence
Multi-document synthesis6Retrieving one facet of a many-part answer
Comparison and conditionals4Negation and qualifiers, where embeddings are weakest
Ambiguous or under-specified3Whether it asks instead of guessing
Unanswerable5Whether it refuses — nothing else measures this

Write the unanswerable ones deliberately: plausible questions in your domain whose answers are genuinely absent from the corpus. If your bot answers those five, it will hallucinate for users, and no other test will tell you.

Measure three things separately

  • Recall@5 — did the gold document IDs appear in the five chunks sent to the model?
  • Faithfulness — split each answer into atomic claims and check each against the context. Four supported claims out of five is 0.80.
  • End-to-end correctness — does the answer match the gold answer, judged by a model with the gold answer in front of it.

These three relate multiplicatively: end-to-end is roughly recall multiplied by faithfulness. That relationship tells you where your headroom is, which is the whole reason for measuring separately.

Optimise by ablation, one change at a time

Change one thing, re-run, record. A representative run looks like this:

ConfigurationRecall@5FaithfulnessEnd-to-end
Baseline: 800-token chunks, no overlap, dense only, k=50.630.710.47
+ 400-token chunks with 15% overlap0.740.760.58
+ hybrid BM25 and dense retrieval0.830.770.66
+ cross-encoder re-rank, fetch 30 → 50.910.790.73
+ grounded prompt, citations, temperature 00.910.940.85

Read the last row carefully. It changed nothing about retrieval — recall stayed at 0.91 — and it was worth twelve points of end-to-end accuracy, more than any single retrieval change in the table. It also took ten minutes.

That is why the project insists you measure retrieval and generation separately. Teams work on retrieval because retrieval feels like the interesting engineering, and the arithmetic frequently says otherwise.

Documentation and packaging — 30 minutes

Organise the code so that each stage can be run and inspected independently:

Text
project/  ingest.py        # load, validate, chunk, embed, index  retrieve.py      # hybrid search + re-ranking  chat.py          # prompt, chain, interface  evaluate.py      # runs the test set, prints the decomposed report  data/            # source documents  eval/questions.json  README.md

The README needs six things: what domain and why, corpus statistics, the pipeline in a paragraph, your final metrics, your ablation table, and the limitations you know about. That ablation table is the most valuable thing in the repository — it is evidence you measured rather than guessed, and it is what an assessor or an interviewer will read first.

Deliverables and marking

ComponentWeightWhat earns full marks
Document preparation15%Domain justified with the control test; 30+ documents; validation report; metadata attached at ingestion
Retriever implementation20%Chunking choice justified by measurement; hybrid search; re-ranking; recall@5 ≥ 0.85
Generator integration20%Grounded prompt with an explicit refusal path; working citations; query condensation; usable interface
Evaluation and optimisation25%30+ stratified questions including unanswerable ones; retrieval and generation scored separately; ablation table with at least four configurations
Documentation and packaging20%Runs from a clean checkout; README covers all six points; separated modules; limitations stated honestly

Stated limitations are marked, not penalised. "Recall drops to 0.6 on questions requiring three or more documents, because our chunking splits multi-part procedures" is a stronger answer than silence, and it is what distinguishes someone who measured from someone who hoped.

If you finish early

  • Multi-hop retrieval. Decompose "how does our refund policy compare to our cancellation policy?" into two retrievals and synthesise. Watch what it does to latency.
  • Semantic caching. Cache answers by query embedding. Then measure your false-hit rate — you will find it is higher than you expect, especially on questions containing numbers or negation.
  • Confidence gating. Combine the top re-rank score with the margin between first and fifth. Refuse below a threshold, and check whether that lifts your refusal accuracy on the five unanswerable questions.
  • Conversation memory with persistent per-thread state, so a restarted process resumes mid-conversation.

Where these projects go wrong

SymptomCauseFix
Great demo, poor evaluation scoresDemo questions were unconsciously written to match your chunkingHave someone else write the test questions
Everything scores 0.95, feels too easyThe model already knows the domainRun the control test with retrieval disabled
Recall stuck below 0.7 whatever you tryAnswers split across chunk boundaries; no chunk contains one wholeSmaller chunks with more overlap; split on headings first
Specific codes and part numbers never foundDense-only retrieval; embeddings have no meaning for such tokensAdd BM25 and merge the result lists
Right documents retrieved, wrong answersWeak prompt with no grounding constraint and no refusal pathConstrain to context, forbid inference, define the refusal sentence, temperature 0
Answers all five unanswerable questionsNo refusal path, or a confidence threshold that never firesExplicit refusal instruction plus a score-based gate
Multi-turn conversation degrades after turn twoPronouns reach the retriever unresolvedCondense every follow-up into a standalone question
Re-ranking changes nothingFetching only 5 candidates — nothing to reorderFetch 30-50, re-rank down to 5

What this project is actually testing

Not whether you can wire five libraries together. That part is an afternoon, and the frameworks do most of it.

It is testing whether you can tell which stage is failing. A working RAG system is a chain of five or six components, each of which can degrade quietly, and the observable symptom — "the bot said something wrong" — is identical whatever the cause. Someone who can look at recall@5 of 0.91 and faithfulness of 0.71 and immediately know the retriever is fine and the prompt is the problem will fix it in ten minutes. Someone with one aggregate accuracy number will spend a week changing embedding models.

So the artefacts to be proud of when you finish are not the chatbot. They are the stratified test set with unanswerable questions in it, the ablation table showing what each change was worth, and the honest paragraph about what still does not work. Those three things are what a production RAG team does every week, and the chatbot is just the thing they happen to be doing it to.