Course Content
RAG Systems
12 sections · 66 lessons
How do you reduce latency in a RAG system?
What you need to know
A latency budget, step by step
A typical request for a bank FAQ bot, with illustrative timings:
| Step | Time | Notes |
|---|---|---|
| Query rewrite (small LLM) | 350 ms | Turns "and for Gold?" into a full question |
| Embed the query | 40 ms | One short text |
| Dense + keyword search | 80 ms | Run in parallel |
| Rerank 40 candidates | 250 ms | Cross-encoder |
| Model prefill (reads 2,000 tokens) | 400 ms | Part of TTFT |
| Generate 250 tokens at 15 ms each | 3,750 ms | The biggest cost |
| Total | about 4.9 s | TTFT is about 1.1 s with streaming |
Two lessons from the arithmetic. Generation dominates, so output length matters: asking for "at most 4 sentences" instead of a free essay can cut seconds. And without streaming the user stares at a spinner for 4.9 seconds; with streaming, they read from about 1.1 seconds.
The levers, in order of payoff
- Stream the response. Same total time, much better experience.
- Shorter output. Every output token costs about the same time; fewer tokens, faster answer.
- Parallelise independent steps.
1import asyncio23async def retrieve(question, user_id):4 dense, keyword, groups = await asyncio.gather(5 dense_search(question), keyword_search(question), load_user_acl(user_id))6 return dense, keyword, groupsWith stand-in functions taking 80, 50 and 30 ms, this finishes in about 80 ms instead of 160 ms, because asyncio.gather runs all three at once and waits for the slowest.
- Fewer, better chunks. Reranking down to 4 chunks instead of 20 cuts prefill time and cost.
- Right-size models. A small model for query rewriting and routing; a large one only for the final answer. Skip the rewrite step for first-turn questions that are already standalone.
- Cache. Embeddings, retrieval results, and whole answers for repeated questions (see the caching lesson). Provider prompt caching speeds up a long, fixed system prompt.
- Infrastructure. Keep the vector index in memory, put the app and the store in the same region, and tune the ANN search setting (for HNSW, lower
ef_searchis faster with slightly lower recall).
Watch the tail
Report p50, p95 and p99. A p50 of 2 seconds with a p99 of 15 seconds means one in a hundred users waits 15 seconds. Tails usually come from retries, cold caches, rate limits, or unusually long answers.
A real-life example
An e-commerce product Q&A widget has a p95 of 7.8 seconds, and the team plans to switch vector databases. The trace breakdown shows the vector search takes 60 ms. The real costs are a query-rewrite call on every question (600 ms), a reranker scoring 150 candidates (1.2 s), and answers averaging 380 tokens.
The fixes: skip the rewrite for single-turn questions, cut rerank candidates to 40, add "answer in at most 3 sentences", and stream. The p95 falls to about 3.5 seconds, and time to first token to about 0.9 seconds. The database switch, which would have saved a few tens of milliseconds, is cancelled.
Follow-up questions to expect
- "Does a smaller embedding model make it faster?" — Slightly, but query embedding is already tens of milliseconds. It matters more at ingestion time and for index memory.
- "How do you speed up the reranker?" — Fewer candidates, a smaller cross-encoder, running it on a GPU, or skipping it when the top retrieval score is already clearly above the rest.
- "What about agentic RAG?" — Each extra loop adds a full model call, often a second or more. Cap the steps and route simple questions to the single-pass path.