RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the step-by-step workflow of a complete RAG pipeline?


The request path for one follow-up questionRewrite withchat historyHybridsearch withfilters, 30 hitsCross-encoderrerank, keep 4Prompt withnumbered sourcesStream answerwith citationsLog query,chunk IDs,scoresLLM calls dominate the latency; search takes tens of milliseconds.
Rewriting and reranking are optional in a demo; filters, the not-found path and logging are not optional in production.

What you need to know

Offline: build the index

  1. Load — read each source into text plus metadata (source, page, version, access groups).
  2. Clean and parse — remove navigation, headers, footers and boilerplate; keep tables and code blocks whole.
  3. Split — make chunks of roughly 200 to 800 tokens, following headings and paragraphs where possible.
  4. Embed and store — embed each chunk, write vector, text and metadata under a stable ID; also add the text to a keyword (BM25) index.

Online: answer a question

  1. Rewrite the query — turn "and the second one?" into a standalone question using chat history.
  2. Retrieve wide — run dense and BM25 search with metadata filters (tenant, permissions, current version), fuse the lists, keep about 30 candidates.
  3. Rerank — score each candidate against the question with a cross-encoder and keep the top 3 to 5.
  4. Build the prompt — instructions, numbered source blocks with labels, then the question.
  5. Generate — stream the answer; require a citation like [2] after each claim.
  6. Log — store the query, chunk IDs, scores, prompt version and answer for tracing and evaluation.

Where the time goes

A typical budget for one request, with made-up but realistic numbers for a mid-size system:

StepTypical time
Query rewrite (small model)150–400 ms
Embed query10–50 ms
Hybrid search10–50 ms
Rerank 30 candidates50–300 ms
Time to first token of the answer300–1,000 ms

The LLM calls dominate. That is why teams skip the rewrite step when there is no chat history, and why they stream the answer so the user sees words early.

Which steps are optional

  • Query rewriting matters only for conversations or messy queries.
  • Hybrid search matters as soon as users type codes, product names or IDs.
  • Reranking matters when the right chunk is often in the top 30 but not the top 5.
  • Logging is never optional in production. Without it you cannot debug a single complaint.

A real-life example

An e-commerce product Q&A feature. A shopper on a phone listing asks, in chat: "Does it support fast charging?" and then "What about the charger in the box?"

  1. The rewrite step turns the second message into "Is a fast charger included in the box for product 88213?".
  2. Hybrid search, filtered to product_id = 88213, returns 30 candidates from the spec sheet, the seller's Q&A and reviews. BM25 finds the exact line "In the box: handset, 33W charger, USB-C cable".
  3. The reranker moves that line to the top.
  4. The prompt holds 4 numbered sources and the question.
  5. The answer: "Yes, a 33W charger is included in the box [1]." The shopper sees the cited spec line under the answer.
  6. The log records all of it, so when a seller disputes an answer, support can replay exactly what the model saw.

Follow-up questions to expect

  • "Where do permission filters go?" — Inside the retrieval step, as a filter on the search itself, never as an instruction to the model.
  • "What if retrieval returns nothing useful?" — If the reranker's top score is below a tuned threshold, reply that the documents do not cover it, or hand off to a human, rather than letting the model guess.
  • "Would you use an agent for this?" — Only if questions need several searches or tools. A fixed pipeline is faster, cheaper and easier to test.