Course Content
Advanced RAG
3 sections · 38 lessons
What is the cold start problem in RAG, and how can it be solved early on?
What you need to know
Two different cold starts
- No content. A new product, a new market, a knowledge base nobody has written. Retrieval quality cannot exceed coverage.
- No traffic. You have documents but no real questions, clicks or ratings. You cannot tune thresholds, compare chunkers or fine-tune anything with evidence.
Seed the content
Ingest whatever exists, even if it is messy: product docs, support macros, resolved tickets, release notes, wiki pages, recorded training sessions (transcribed). You can improve chunking later; you cannot retrieve a document you never ingested. Mark low-trust sources (old Slack threads) with metadata so they rank below official documents.
Build an evaluation set with zero traffic
- Sample 200–500 chunks, stratified across document types.
- Generate — an LLM writes one or two realistic questions each chunk answers, in the voice of real users (short, informal, sometimes misspelled or in the users' languages).
- Filter — a human spot-checks a sample and drops questions that are trivial copies of the chunk's wording or unanswerable.
- Measure retrieval — each (question → source chunk) pair is ground truth for recall@k and MRR.
- Measure answers — for a subset, write reference answers and score faithfulness and correctness.
1def recall_at_k(results, gold, k):2 return sum(gold[q] in results[q][:k] for q in gold) / len(gold)34def mrr(results, gold):5 total = 0.06 for q, g in gold.items():7 if g in results[q]:8 total += 1 / (results[q].index(g) + 1)9 return total / len(gold)1011gold = {"q1": "c17", "q2": "c03", "q3": "c88", "q4": "c41"}12results = {"q1": ["c17", "c02", "c09"], "q2": ["c11", "c03", "c40"],13 "q3": ["c05", "c06", "c07"], "q4": ["c41", "c40", "c39"]}14print(recall_at_k(results, gold, 1), recall_at_k(results, gold, 3)) # 0.5 0.7515print(round(mrr(results, gold), 3)) # 0.625Recall@k asks "was the right chunk anywhere in the top k?"; MRR (mean reciprocal rank) rewards putting it near the top. With these numbers you can compare chunk sizes, embedding models and rerankers on evidence instead of opinion.
Caveats about synthetic data and LLM judges
- Synthetic questions are easier than real ones: they reuse the chunk's own words. Replace them with real queries as soon as traffic arrives.
- LLM judges have known biases: they prefer longer answers, can favour their own model family's style, and in pairwise comparisons are swayed by position (which answer is shown first). Calibrate the judge on 50–100 human-labelled examples, swap positions in pairwise tests, and use a judge from a different model family where possible.
Ship defaults, not tuning
Hybrid search, structure-aware chunks of a few hundred tokens with small overlap, a cross-encoder reranker, top-5. These work reasonably on most corpora without data. Defer what needs traffic: semantic caching, learned routing, embedding fine-tuning.
Make gaps visible
If the best reranker score is below a floor, answer "I don't have information on that" and log the query. The ranked list of unanswered questions is the most valuable content roadmap you will get. Instrument thumbs up/down, citation clicks, and "user rephrased immediately" as an implicit negative from day one.
A real-life example
A company launches an engineering-wiki assistant for 800 engineers. There is no query log — the assistant is new. The team:
- ingests 14,000 wiki pages, 2,000 runbooks and the last year of resolved incident tickets;
- generates 400 synthetic questions from sampled chunks, and asks 10 senior engineers to write 100 more "real" ones; the human-written set turns out much harder;
- compares two chunk sizes and two embedding models on recall@5 — the choice they would otherwise have made by opinion flips once measured;
- ships with a "no good source" floor.
In the first month, the unanswered-query log shows 60 questions about the new internal developer platform, which had no documentation at all. The platform team writes six pages, and that category of question moves from "no answer" to answered. The feedback data then lets the team replace the synthetic set with 600 real, labelled questions.
Follow-up questions to expect
- "How do you stop synthetic questions from being too easy?" — Prompt for paraphrase and informal wording, forbid copying phrases from the chunk, add multi-chunk questions, and mix in human-written questions.
- "How do you validate an LLM judge?" — Compare its scores with human labels on a sample and measure agreement; re-check whenever you change the judge model or prompt.
- "What would you tune first once traffic arrives?" — The no-answer threshold and the retrieval settings, using real failed queries; then caching and routing.