RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you reduce latency in a RAG system?


What you need to know

A latency budget, step by step

A typical request for a bank FAQ bot, with illustrative timings:

StepTimeNotes
Query rewrite (small LLM)350 msTurns "and for Gold?" into a full question
Embed the query40 msOne short text
Dense + keyword search80 msRun in parallel
Rerank 40 candidates250 msCross-encoder
Model prefill (reads 2,000 tokens)400 msPart of TTFT
Generate 250 tokens at 15 ms each3,750 msThe biggest cost
Totalabout 4.9 sTTFT is about 1.1 s with streaming

Two lessons from the arithmetic. Generation dominates, so output length matters: asking for "at most 4 sentences" instead of a free essay can cut seconds. And without streaming the user stares at a spinner for 4.9 seconds; with streaming, they read from about 1.1 seconds.

The levers, in order of payoff

  1. Stream the response. Same total time, much better experience.
  2. Shorter output. Every output token costs about the same time; fewer tokens, faster answer.
  3. Parallelise independent steps.
Python
import asyncioasync def retrieve(question, user_id):    dense, keyword, groups = await asyncio.gather(        dense_search(question), keyword_search(question), load_user_acl(user_id))    return dense, keyword, groups

With stand-in functions taking 80, 50 and 30 ms, this finishes in about 80 ms instead of 160 ms, because asyncio.gather runs all three at once and waits for the slowest.

  1. Fewer, better chunks. Reranking down to 4 chunks instead of 20 cuts prefill time and cost.
  2. Right-size models. A small model for query rewriting and routing; a large one only for the final answer. Skip the rewrite step for first-turn questions that are already standalone.
  3. Cache. Embeddings, retrieval results, and whole answers for repeated questions (see the caching lesson). Provider prompt caching speeds up a long, fixed system prompt.
  4. Infrastructure. Keep the vector index in memory, put the app and the store in the same region, and tune the ANN search setting (for HNSW, lower ef_search is faster with slightly lower recall).

Watch the tail

Report p50, p95 and p99. A p50 of 2 seconds with a p99 of 15 seconds means one in a hundred users waits 15 seconds. Tails usually come from retries, cold caches, rate limits, or unusually long answers.

A real-life example

An e-commerce product Q&A widget has a p95 of 7.8 seconds, and the team plans to switch vector databases. The trace breakdown shows the vector search takes 60 ms. The real costs are a query-rewrite call on every question (600 ms), a reranker scoring 150 candidates (1.2 s), and answers averaging 380 tokens.

The fixes: skip the rewrite for single-turn questions, cut rerank candidates to 40, add "answer in at most 3 sentences", and stream. The p95 falls to about 3.5 seconds, and time to first token to about 0.9 seconds. The database switch, which would have saved a few tens of milliseconds, is cancelled.

Follow-up questions to expect

  • "Does a smaller embedding model make it faster?" — Slightly, but query embedding is already tens of milliseconds. It matters more at ingestion time and for index memory.
  • "How do you speed up the reranker?" — Fewer candidates, a smaller cross-encoder, running it on a GPU, or skipping it when the top retrieval score is already clearly above the rest.
  • "What about agentic RAG?" — Each extra loop adds a full model call, often a second or more. Cap the steps and route simple questions to the single-pass path.