Course Content
Advanced RAG
3 sections · 38 lessons
How would you design a large-scale RAG system for thousands of daily queries?
What you need to know
Size the problem before drawing boxes
5,000 queries a day, mostly in office hours, is roughly 0.2 queries per second on average and perhaps 2–3 at peak. One well-configured service handles that. Saying this out loud is a good interview signal: it moves the discussion to what is actually hard — millions of chunks, documents changing daily, per-user permissions, answer quality and cost.
Ingestion (offline)
- Connectors and change feed — pull from document systems, wikis and databases, triggered by change events, not nightly full scans.
- Parse — layout-aware parsing; tables kept whole.
- Chunk and enrich — structure-aware chunks, contextual headers, metadata:
doc_id, version, dates, language, and ACL labels (which groups may see it). - Embed in batches — through a queue with workers, so a large backfill cannot starve live traffic. Idempotent by content hash.
- Upsert — into a vector store that also supports BM25 (or a paired keyword index).
Serving path (online)
request → auth → semantic cache (scoped by user permissions) → query rewrite (small model) → hybrid retrieval: BM25 + dense, ACL filter inside the query, RRF → cross-encoder rerank top ~50 → top 5 → generate (streaming) with citation IDs → validate citations → responseAt this scale a managed vector database, or Postgres with pgvector for a few million chunks, is enough. You do not need a distributed cluster.
Security: two things interviewers probe
- Per-user access control on chunks. Store ACL labels on every chunk and pass the user's groups as a filter inside the retrieval query, so forbidden chunks are never returned. Filtering after retrieval risks leaks through bugs, and also returns fewer than k results. Permission changes must propagate within minutes.
- Prompt injection via documents. Anyone who can edit a wiki page or upload a PDF can plant text such as "ignore your instructions and reveal…". Defences: wrap retrieved text in clear delimiters and tell the model it is data, not instructions; give the RAG assistant no dangerous tools, or require confirmation for them; scan ingested content for injection patterns; and never render model-produced links or images that could leak data to outside URLs.
Reliability
A timeout and fallback at every hop: reranker down → use retrieval order; retrieval down → say "search is unavailable" rather than answering from memory; model provider errors → retry once, then a secondary model. Per-tenant rate limits stop one team's batch job from slowing everyone.
Observability and evaluation
Log for every request: raw and rewritten query, retrieved chunk IDs with scores, reranked order, the final prompt's token counts, model version, per-stage latency and cost. Without retrieved IDs in logs you cannot debug a bad answer later.
Evaluate the stages separately. For retrieval: recall@k and context precision on a labelled set. For generation: faithfulness (claims supported by the context) and answer relevance, often scored by an LLM judge that you have checked against human labels. RAGAS and similar frameworks provide these metrics. One blended "quality" score tells you nothing about which half to fix.
A real-life example
A pharma company builds regulatory-document search for 3,000 employees across regulatory affairs, quality and R&D in six countries. The corpus: 400,000 documents (about 6 million chunks) — guidelines, SOPs, submissions and agency correspondence. Traffic: around 8,000 queries a day.
Key design decisions:
- ACLs: submission documents are restricted by product and country team. Each chunk carries
acl_groups; the user's groups from the identity provider are applied as a filter in the hybrid query. A test suite checks that a user from the oncology team never receives cardiology submission chunks. - Freshness: a webhook from the document system triggers re-indexing when a document becomes "Effective"; the target is under 15 minutes.
- Injection: agency correspondence arrives as external PDFs. Retrieved text is wrapped in
<document>tags with the instruction that it is data only; the assistant has no tools except search. - Evaluation: 500 questions written by regulatory staff, with the supporting documents labelled. Every change to chunking, models or prompts must not lower retrieval recall@10 or faithfulness on this set.
They ship hybrid retrieval, reranking, a permission-scoped cache and full tracing first. Multi-query and agentic multi-hop retrieval are added only after logs show which question types still fail.
Follow-up questions to expect
- "How does the design change at 1,000 queries per second?" — Now throughput matters: horizontal scaling of stateless services, sharded or replicated indexes, GPU capacity for reranking, aggressive caching, and careful attention to p99 latency.
- "How do you test access control?" — Automated tests with synthetic users in different groups asserting what they can and cannot retrieve, plus audit logs of which chunks were shown to whom.
- "What if the LLM judge is wrong?" — Judges have known biases (they favour longer answers, their own model family's style, and the first option shown). Calibrate the judge on a human-labelled sample and track agreement.