Course Content
RAG Systems
12 sections · 66 lessons
How does multi-step retrieval improve answer quality?
What you need to know
Why one search is not enough
"How did the refund policy change after the 2025 audit?" needs two facts from two documents: what the audit found, and what the current policy says. Their wording is different, so one embedding of the whole question sits between them and may retrieve neither well.
The main patterns
| Pattern | What it does | Example |
|---|---|---|
| Query rewriting | Makes a follow-up standalone | "and for Gold?" becomes "Gold card annual fee" |
| Decomposition | Splits into independent sub-questions, searched in parallel | "Compare the return policy for shoes and watches" becomes two searches |
| Multi-hop | Each search depends on the previous answer | "Who approves leave for my manager's team?" finds the manager, then the approval rule |
| Step-back | Searches a broader question first | "Why was my claim for a 4-star hotel rejected?" first retrieves the travel policy's hotel rules |
| HyDE | Writes a hypothetical answer and searches with it | Useful when questions are short and documents are long |
Query rewriting is cheap and helps nearly every chat-based RAG system. The others are for specific question types.
Latency arithmetic
With illustrative numbers: a planning call 0.6 s, each hop 0.3 s retrieval plus 0.6 s for the model to read results and write the next query, and a final answer 3 s.
single pass: 0.3 + 3.0 = 3.3 sdecomposed (3, parallel): 0.6 + 0.3 + 3.0 = 3.9 smulti-hop (3 hops): 0.6 + 3 x (0.3 + 0.6) + 3.0 = 6.3 sDecomposition runs searches at the same time, so it adds little. Multi-hop is sequential, and every hop adds almost a second. Errors add up the same way: if each hop picks the right thing 90% of the time, three hops are all right only about 0.9 × 0.9 × 0.9 ≈ 73% of the time.
GraphRAG for "connect the dots" questions
Some questions ask about the whole corpus: "What are the main complaints about our home loans this quarter?" No chunk holds that answer. GraphRAG (published by Microsoft in 2024) uses an LLM at index time to extract entities and relationships into a knowledge graph, groups related entities into communities, and writes a summary for each. Global questions are answered from those summaries. It is powerful for themes and relationships across many documents, but indexing needs many LLM calls over the whole corpus, so it is expensive to build and to refresh. Lighter variants, such as Microsoft's LazyGraphRAG, defer most of that work to query time. Use it for sensemaking over a stable corpus, not for FAQ lookups.
A real-life example
An e-commerce product Q&A assistant gets: "Will the charger that comes with the Nova 12 phone fast-charge my Aero laptop?" The answer requires: which charger ships with the Nova 12 (from the phone's spec), what that charger outputs, and what the Aero laptop accepts.
Single-pass retrieval returns the phone's spec page and says "yes, 33 W fast charging", which is about the phone, not the laptop. The multi-hop path first retrieves "Nova 12 box contents" (33 W USB-C charger), then "Aero laptop charging input" (needs 65 W USB-C Power Delivery), and answers: "It may charge the laptop slowly, or not at all while it is in use, because the 33 W charger is below the laptop's 65 W requirement [1][2]." The assistant only takes this path for questions naming two or more products, about 8% of traffic.
Follow-up questions to expect
- "How do you decide if a question needs multiple steps?" — A cheap classifier or rules (two product names, "compare", "before and after"), checked against the golden set.
- "How do you avoid repeating the same chunks across hops?" — Keep a set of chunk ids already seen and drop duplicates before adding to context.
- "When would you choose GraphRAG over normal RAG?" — When important questions are about themes or relationships across many documents and the corpus changes slowly enough to justify the indexing cost.