RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you design a RAG system for production use?


What you need to know

In a design interview, the architecture matters less than showing why each part exists. Start with questions.

1. Requirements to ask for

QuestionWhy it matters
How many documents, how often do they change?Index size, ingestion design, freshness
Queries per day, peak per second?Scaling, cost
Latency target?Streaming, model size, reranker
Who may see what?Access control design, tenant isolation
Cost of a wrong answer?Guardrails, human review, refusals
Where must data stay?Model hosting, region, self-hosting

2. The architecture

  1. Ingestion service — connectors, parsing (with OCR and table handling), cleaning, chunking, metadata including permissions, embedding with a cache, sync by content hash, and nightly reconciliation.
  2. Index — vector plus keyword index, permission and version fields on every chunk, versioned collections for blue/green changes.
  3. Query service — optional query rewrite, hybrid search with permission filters inside the store, reranking to 3 to 6 chunks, relevance gate.
  4. Generation — grounded prompt with labelled sources, required citations, streaming, citation check before returning.
  5. Evaluation — golden set in CI blocking bad releases, sampled faithfulness scoring in production.
  6. Operations — tracing, dashboards (recall on golden set, faithfulness, p95 latency, cost per query, freshness lag, cache hit rate), alerts, audit logs.

3. Put numbers on it

For an HR assistant with 20,000 employees: about 5,000 questions a day, peaking around 3 per second at 10 a.m. With 2,500 input tokens and 300 output tokens per answer, at assumed prices of $3 per million input and $15 per million output tokens:

Text
input:  2,500 x $3  / 1,000,000 = $0.0075output:   300 x $15 / 1,000,000 = $0.0045per query                        = $0.012per day   (x 5,000)              = $60

Retrieval infrastructure and the reranker add to this, but the model call usually dominates. This arithmetic tells the interviewer you can budget, and shows where savings would come from (shorter output, fewer chunks, a smaller model for easy questions).

4. What to build first

Version 1: ingestion with deterministic ids, hybrid retrieval, permission filters, citations, the "not found" gate, tracing and a 100-question golden set. Add reranking when the golden set shows ranking problems, caching when traffic justifies it, and agentic steps only for question types that single-pass retrieval fails on.

A real-life example

A hospital builds guideline search for 3,000 clinical staff. Requirements that shape the design: wrong answers can hurt patients; guidelines change weekly; some documents are restricted to specific departments; patient data must not leave the country.

So the design uses a model hosted in-country; department permissions stored on every chunk and filtered inside the store; a strict relevance gate with a clear "not found, see the full guideline" fallback; every answer showing the guideline name, version and section; a golden set of 300 questions written by clinicians, including 50 with no answer in the corpus; and a weekly review of low-faithfulness answers by a clinical informatics nurse. Caching is left out of version 1 entirely, because guidelines change and questions are specific. The design is chosen by the risk, not by what is fashionable.

Follow-up questions to expect

  • "How would you scale to 100 times the traffic?" — The query service is stateless and scales horizontally; the vector store scales with replicas and shards; model calls scale with provider rate limits, so you add caching, smaller models for simple questions, and queueing.
  • "How do you roll out a change safely?" — Golden set in CI, then a shadow run or a small percentage of traffic, compare metrics, then full rollout.
  • "What would you monitor first?" — Empty-retrieval rate, faithfulness on a sample, p95 latency, cost per query, and freshness lag.