Course Content
RAG Systems
12 sections · 66 lessons
How do you design a RAG system for production use?
What you need to know
In a design interview, the architecture matters less than showing why each part exists. Start with questions.
1. Requirements to ask for
| Question | Why it matters |
|---|---|
| How many documents, how often do they change? | Index size, ingestion design, freshness |
| Queries per day, peak per second? | Scaling, cost |
| Latency target? | Streaming, model size, reranker |
| Who may see what? | Access control design, tenant isolation |
| Cost of a wrong answer? | Guardrails, human review, refusals |
| Where must data stay? | Model hosting, region, self-hosting |
2. The architecture
- Ingestion service — connectors, parsing (with OCR and table handling), cleaning, chunking, metadata including permissions, embedding with a cache, sync by content hash, and nightly reconciliation.
- Index — vector plus keyword index, permission and version fields on every chunk, versioned collections for blue/green changes.
- Query service — optional query rewrite, hybrid search with permission filters inside the store, reranking to 3 to 6 chunks, relevance gate.
- Generation — grounded prompt with labelled sources, required citations, streaming, citation check before returning.
- Evaluation — golden set in CI blocking bad releases, sampled faithfulness scoring in production.
- Operations — tracing, dashboards (recall on golden set, faithfulness, p95 latency, cost per query, freshness lag, cache hit rate), alerts, audit logs.
3. Put numbers on it
For an HR assistant with 20,000 employees: about 5,000 questions a day, peaking around 3 per second at 10 a.m. With 2,500 input tokens and 300 output tokens per answer, at assumed prices of $3 per million input and $15 per million output tokens:
input: 2,500 x $3 / 1,000,000 = $0.0075output: 300 x $15 / 1,000,000 = $0.0045per query = $0.012per day (x 5,000) = $60Retrieval infrastructure and the reranker add to this, but the model call usually dominates. This arithmetic tells the interviewer you can budget, and shows where savings would come from (shorter output, fewer chunks, a smaller model for easy questions).
4. What to build first
Version 1: ingestion with deterministic ids, hybrid retrieval, permission filters, citations, the "not found" gate, tracing and a 100-question golden set. Add reranking when the golden set shows ranking problems, caching when traffic justifies it, and agentic steps only for question types that single-pass retrieval fails on.
A real-life example
A hospital builds guideline search for 3,000 clinical staff. Requirements that shape the design: wrong answers can hurt patients; guidelines change weekly; some documents are restricted to specific departments; patient data must not leave the country.
So the design uses a model hosted in-country; department permissions stored on every chunk and filtered inside the store; a strict relevance gate with a clear "not found, see the full guideline" fallback; every answer showing the guideline name, version and section; a golden set of 300 questions written by clinicians, including 50 with no answer in the corpus; and a weekly review of low-faithfulness answers by a clinical informatics nurse. Caching is left out of version 1 entirely, because guidelines change and questions are specific. The design is chosen by the risk, not by what is fashionable.
Follow-up questions to expect
- "How would you scale to 100 times the traffic?" — The query service is stateless and scales horizontally; the vector store scales with replicas and shards; model calls scale with provider rate limits, so you add caching, smaller models for simple questions, and queueing.
- "How do you roll out a change safely?" — Golden set in CI, then a shadow run or a small percentage of traffic, compare metrics, then full rollout.
- "What would you monitor first?" — Empty-retrieval rate, faithfulness on a sample, p95 latency, cost per query, and freshness lag.