Course Content
Retrieval-Augmented Generation (RAG)
4 sections · 8 lessons
Mini Project: Domain-Specific RAG Chatbot
Almost everyone's first RAG chatbot demos beautifully and then falls over in public.
The pattern is consistent. You load 40 PDFs, split them at 1,000 characters, embed them, wire up a retriever and a prompt, and ask five questions you already know the answers to. All five come back right. You show someone. They ask a sixth question — "does that apply to contractors as well?" — and the bot invents a policy that does not exist, in the same confident tone it used for the five correct answers.
Nothing malfunctioned. The retriever returned its top 5 chunks, none of which mentioned contractors, and the model filled the gap because nothing in the prompt told it not to. The five questions you tested with were the five you had unconsciously written to match your chunking.
This project is the fix. You will build a domain-specific RAG chatbot end to end, and — the part that matters — you will measure it honestly enough to know which stage is failing and what each change actually bought you. Budget about four and a half hours.
What you are building
A question-answering system over a document collection you choose, meeting concrete acceptance criteria rather than "it works":
| Requirement | Target |
|---|---|
| Corpus size | 30-100 documents, 150,000+ words total |
| Retrieval quality | Recall@5 ≥ 0.85 on your test set |
| Faithfulness | ≥ 0.90 — claims supported by retrieved context |
| Refusal behaviour | Correctly declines ≥ 80% of unanswerable questions |
| Citations | Every answer names its source documents |
| Latency | p95 under 5 seconds end to end |
| Evaluation | Scripted, reproducible, reports retrieval and generation separately |
The refusal row is what separates a real system from a demo. A chatbot that answers everything scores perfectly on a test set where everything is answerable, and hallucinates the moment it meets a real user.
Data preparation — 45 minutes
Choose a domain the model does not already know
This decision determines whether the whole project teaches you anything, and most people get it wrong.
Pick "Python documentation" or "the history of the Roman Empire" and your chatbot will answer well with the retriever completely disconnected, because the model already knows the material. You will have built nothing and measured nothing.
Run the control test before you write a line of pipeline code. Take 20 questions from your intended domain and ask a bare language model with no retrieval at all.
| Bare model gets right | Verdict |
|---|---|
| 15-20 of 20 | Wrong domain. The model knows it. Choose again. |
| 6-14 of 20 | Usable, but your measured gains will be muddied. |
| 0-5 of 20 | Correct choice. Every point of accuracy came from your system. |
Domains that reliably pass: your own organisation's internal documentation; a niche regulatory corpus; product manuals for specific equipment; a set of research papers from the last twelve months; the complete rules of a game with a small following; the minutes and bylaws of a local institution.
Collect and validate
Aim for 30-100 documents totalling 150,000 words or more — roughly 200 pages. Below that, retrieval is trivially easy and you will not see the failure modes the project exists to teach.
Validate before indexing, because every extraction problem becomes a retrieval problem you will misdiagnose for an hour:
1from pathlib import Path23def validate(docs):4 report = []5 for d in docs:6 text = d.page_content7 words = len(text.split())8 alpha = sum(c.isalpha() for c in text) / max(len(text), 1)9 report.append({10 "source": d.metadata["source"],11 "words": words,12 "alpha": round(alpha, 2), # < 0.6 suggests bad extraction13 "empty": words < 50,14 "no_space": " " not in text[:200], # PDF ligature failure15 })16 return reportThe two checks that catch the most damage are the alphabetic ratio and the missing-space test. A PDF that extracted as "Refundswithinfourteendays" looks fine in a file listing, embeds to something meaningless, and silently removes that document from your corpus. You will spend an hour blaming the embedding model.
Attach metadata now, not later: source, doc_id, title, section, date, and any category you might want to filter on. Retrofitting metadata means re-embedding everything.
Retriever implementation — 60 minutes
Chunking is the decision that matters most
Most beginner RAG failures are chunking failures wearing a costume. A chunk that splits a policy from its exception produces a retriever that confidently returns half a rule.
| Chunk size | Chunks from 200k words (15% overlap) | Recall behaviour | Failure mode |
|---|---|---|---|
| 200 tokens | ~1,600 | Sharp, focused embeddings | Answers split across chunks; no chunk is complete |
| 400 tokens | ~800 | Best general-purpose choice | Occasional split at section boundaries |
| 800 tokens | ~400 | Whole sections survive intact | Diluted embeddings; one chunk covers three topics |
| 1,500 tokens | ~215 | Nothing gets split | Retrieval becomes near-random; huge context cost |
Start at 400 tokens with 15% overlap. The overlap exists so that a sentence straddling a boundary appears whole in at least one chunk — without it, "…the window is 14 days. | Contractors are excluded…" produces two chunks, neither of which answers "how long do contractors get?"
1from langchain_text_splitters import RecursiveCharacterTextSplitter23splitter = RecursiveCharacterTextSplitter(4 chunk_size=1600, # ~400 tokens5 chunk_overlap=240, # 15%6 separators=["\n## ", "\n### ", "\n\n", "\n", ". ", " "],7)8chunks = splitter.split_documents(docs)The separators list is doing quiet, important work. It tries to break at a heading first, then a paragraph, then a line, then a sentence, and only mid-word as a last resort. Left at the default, you get splits in the middle of sentences and a measurable recall drop.
Index and configure
Embedding 200,000 words is around 270,000 tokens — a few pence and under two minutes. Cost is not a constraint at this scale, so use a good embedding model rather than a cheap one.
Configure the retriever to fetch wide, not narrow. Fetch 30 candidates and send 5 to the generator. Retrieving 5 directly is the most common configuration mistake in beginner projects: it gives the re-ranker nothing to work with and puts a hard ceiling on accuracy equal to your recall@5.
1retriever = store.as_retriever(search_kwargs={"k": 30})23def get_context(query, k=5):4 candidates = retriever.invoke(query)5 scored = reranker.predict([(query, c.page_content) for c in candidates])6 ranked = sorted(zip(scored, candidates), key=lambda x: x[0], reverse=True)7 return [d for _, d in ranked[:k]]Add BM25 keyword search alongside the vector search and merge the two result lists. Every real corpus contains error codes, part numbers, statute references and proper nouns that embeddings handle badly, because those tokens carry no learned meaning. Keyword search finds them exactly.
Generator integration — 60 minutes
The prompt is where hallucination is prevented
Compare two prompts that look equally reasonable.
WEAK:Answer the question using the context below.Context: {context}Question: {question}STRONG:You answer questions using ONLY the provided context.Rules:1. Every factual statement must be supported by the context. Do not infer, generalise, or fill gaps from your own knowledge.2. Cite the source after each claim, as [source: filename].3. If the context does not contain the answer, reply exactly: "I could not find this in the available documents." Then name the closest topics the documents do cover.4. If sources disagree, say so and quote both.Context:{context}Question: {question}Answer:The weak prompt permits everything the strong one forbids. "Using the context" does not say only the context, does not forbid inference, and offers no escape hatch — so when the context is silent, the model's only available behaviour is to guess. In the ablation below, moving from the weak prompt to the strong one with temperature at 0 was worth twelve points of end-to-end accuracy and cost nothing.
Rule 3 matters more than it looks. Give a model an explicit sentence to say when it does not know, and it will use it. Leave the refusal undefined and it will not invent one.
Wire the chain and put a face on it
1def format_context(docs):2 return "\n\n".join(3 f"[source: {d.metadata['source']}]\n{d.page_content}" for d in docs4 )56def answer(question, history=None):7 standalone = condense(history, question) if history else question8 docs = get_context(standalone, k=5)9 prompt = STRONG.format(context=format_context(docs), question=standalone)10 text = llm.invoke(prompt).content # llm built once from config; temperature 0 if the model allows it11 return text, [d.metadata["source"] for d in docs]Two details carry weight. Putting the source marker inside the context block is what makes citation possible — the model cannot cite what it cannot see. And condense rewrites follow-up questions into standalone ones, so "does that apply to contractors?" reaches the retriever as "does the 14-day refund window apply to contractors?". Skip that step and every multi-turn conversation retrieves noise from turn two onwards.
The interface can be a command-line loop or a small Streamlit app. What it must show, either way, is the retrieved sources and their scores next to every answer. You cannot debug what you cannot see, and a bot that displays its evidence is the single most useful debugging tool in the project.
Evaluation and optimisation — 75 minutes
This phase carries the largest share of the marks, because it is the part that turns a demo into engineering.
Build the test set first, and stratify it
Thirty questions minimum. For each, record the question, the gold answer, and the document IDs that contain it — that last field is what lets you score retrieval separately from generation.
| Type | Count | What it exposes |
|---|---|---|
| Single-fact lookup | 12 | Baseline competence |
| Multi-document synthesis | 6 | Retrieving one facet of a many-part answer |
| Comparison and conditionals | 4 | Negation and qualifiers, where embeddings are weakest |
| Ambiguous or under-specified | 3 | Whether it asks instead of guessing |
| Unanswerable | 5 | Whether it refuses — nothing else measures this |
Write the unanswerable ones deliberately: plausible questions in your domain whose answers are genuinely absent from the corpus. If your bot answers those five, it will hallucinate for users, and no other test will tell you.
Measure three things separately
- Recall@5 — did the gold document IDs appear in the five chunks sent to the model?
- Faithfulness — split each answer into atomic claims and check each against the context. Four supported claims out of five is 0.80.
- End-to-end correctness — does the answer match the gold answer, judged by a model with the gold answer in front of it.
These three relate multiplicatively: end-to-end is roughly recall multiplied by faithfulness. That relationship tells you where your headroom is, which is the whole reason for measuring separately.
Optimise by ablation, one change at a time
Change one thing, re-run, record. A representative run looks like this:
| Configuration | Recall@5 | Faithfulness | End-to-end |
|---|---|---|---|
| Baseline: 800-token chunks, no overlap, dense only, k=5 | 0.63 | 0.71 | 0.47 |
| + 400-token chunks with 15% overlap | 0.74 | 0.76 | 0.58 |
| + hybrid BM25 and dense retrieval | 0.83 | 0.77 | 0.66 |
| + cross-encoder re-rank, fetch 30 → 5 | 0.91 | 0.79 | 0.73 |
| + grounded prompt, citations, temperature 0 | 0.91 | 0.94 | 0.85 |
Read the last row carefully. It changed nothing about retrieval — recall stayed at 0.91 — and it was worth twelve points of end-to-end accuracy, more than any single retrieval change in the table. It also took ten minutes.
That is why the project insists you measure retrieval and generation separately. Teams work on retrieval because retrieval feels like the interesting engineering, and the arithmetic frequently says otherwise.
Documentation and packaging — 30 minutes
Organise the code so that each stage can be run and inspected independently:
project/ ingest.py # load, validate, chunk, embed, index retrieve.py # hybrid search + re-ranking chat.py # prompt, chain, interface evaluate.py # runs the test set, prints the decomposed report data/ # source documents eval/questions.json README.mdThe README needs six things: what domain and why, corpus statistics, the pipeline in a paragraph, your final metrics, your ablation table, and the limitations you know about. That ablation table is the most valuable thing in the repository — it is evidence you measured rather than guessed, and it is what an assessor or an interviewer will read first.
Deliverables and marking
| Component | Weight | What earns full marks |
|---|---|---|
| Document preparation | 15% | Domain justified with the control test; 30+ documents; validation report; metadata attached at ingestion |
| Retriever implementation | 20% | Chunking choice justified by measurement; hybrid search; re-ranking; recall@5 ≥ 0.85 |
| Generator integration | 20% | Grounded prompt with an explicit refusal path; working citations; query condensation; usable interface |
| Evaluation and optimisation | 25% | 30+ stratified questions including unanswerable ones; retrieval and generation scored separately; ablation table with at least four configurations |
| Documentation and packaging | 20% | Runs from a clean checkout; README covers all six points; separated modules; limitations stated honestly |
Stated limitations are marked, not penalised. "Recall drops to 0.6 on questions requiring three or more documents, because our chunking splits multi-part procedures" is a stronger answer than silence, and it is what distinguishes someone who measured from someone who hoped.
If you finish early
- Multi-hop retrieval. Decompose "how does our refund policy compare to our cancellation policy?" into two retrievals and synthesise. Watch what it does to latency.
- Semantic caching. Cache answers by query embedding. Then measure your false-hit rate — you will find it is higher than you expect, especially on questions containing numbers or negation.
- Confidence gating. Combine the top re-rank score with the margin between first and fifth. Refuse below a threshold, and check whether that lifts your refusal accuracy on the five unanswerable questions.
- Conversation memory with persistent per-thread state, so a restarted process resumes mid-conversation.
Where these projects go wrong
| Symptom | Cause | Fix |
|---|---|---|
| Great demo, poor evaluation scores | Demo questions were unconsciously written to match your chunking | Have someone else write the test questions |
| Everything scores 0.95, feels too easy | The model already knows the domain | Run the control test with retrieval disabled |
| Recall stuck below 0.7 whatever you try | Answers split across chunk boundaries; no chunk contains one whole | Smaller chunks with more overlap; split on headings first |
| Specific codes and part numbers never found | Dense-only retrieval; embeddings have no meaning for such tokens | Add BM25 and merge the result lists |
| Right documents retrieved, wrong answers | Weak prompt with no grounding constraint and no refusal path | Constrain to context, forbid inference, define the refusal sentence, temperature 0 |
| Answers all five unanswerable questions | No refusal path, or a confidence threshold that never fires | Explicit refusal instruction plus a score-based gate |
| Multi-turn conversation degrades after turn two | Pronouns reach the retriever unresolved | Condense every follow-up into a standalone question |
| Re-ranking changes nothing | Fetching only 5 candidates — nothing to reorder | Fetch 30-50, re-rank down to 5 |
What this project is actually testing
Not whether you can wire five libraries together. That part is an afternoon, and the frameworks do most of it.
It is testing whether you can tell which stage is failing. A working RAG system is a chain of five or six components, each of which can degrade quietly, and the observable symptom — "the bot said something wrong" — is identical whatever the cause. Someone who can look at recall@5 of 0.91 and faithfulness of 0.71 and immediately know the retriever is fine and the prompt is the problem will fix it in ten minutes. Someone with one aggregate accuracy number will spend a week changing embedding models.
So the artefacts to be proud of when you finish are not the chatbot. They are the stratified test set with unanswerable questions in it, the ablation table showing what each change was worth, and the honest paragraph about what still does not work. Those three things are what a production RAG team does every week, and the chatbot is just the thing they happen to be doing it to.