Course Content
Advanced RAG
3 sections · 38 lessons
How does iterative query refinement improve answers?
What you need to know
Why one pass is not enough
Some questions cannot be answered with one search, because the second search depends on the first result. "Who owns the service that still calls the deprecated v1 ledger API?" needs two hops: first find the service, then find its owner. No single chunk contains both facts.
The loop
- Retrieve with the current query.
- Grade — a small model or reranker threshold checks: is the question answered? What is missing?
- Refine — write a new query aimed at the missing piece.
- Accumulate — keep new documents, drop duplicates.
- Stop — when answered, after 2–3 rounds, or when a round adds nothing new. Then generate.
Kinds of refinement
- Rewriting — fix vocabulary, expand acronyms, turn a follow-up ("and for enterprise?") into a standalone question.
- Step-back prompting — ask a broader question first ("what rules govern data retention for logs?") to fetch the principle, then combine it with the specific query. Useful when the literal question is too narrow to match anything.
- Self-query — have the model turn "incidents in the payments repo since March" into a filter
repo = "payments" AND date >= "2026-03-01"plus a semantic query, so the index does the filtering. - Decomposition — split a compound question into sub-questions and retrieve for each.
- Multi-hop — use an entity found in round 1 (a service name, a drug, a counterparty) as the query for round 2. Methods like IRCoT interleave these retrieval steps with chain-of-thought reasoning.
Agentic RAG
When an LLM runs this loop with retrieval as a tool, deciding itself what to search next and when to stop, it is called agentic RAG. Modern "deep research" features are the long version of this loop, running dozens of searches. For a product feature you want a much tighter version: a few rounds, a time budget, and clear logs of every query the agent issued.
Termination and cost
Every round adds a grading call, a rewrite call and a search — often 1–2 seconds. Without limits, a question the corpus cannot answer loops forever, reformulating each time. So set a maximum number of rounds, a total time budget, and a "no new documents" stop rule. When the loop ends without an answer, say so honestly.
A real-life example
A software company's engineering-wiki assistant gets this question from a new engineer:
Who do I talk to about the service that still uses the v1 ledger API?- Round 1. Query: "service using v1 ledger API". It retrieves the ledger team's deprecation page, which says "remaining v1 consumers:
refund-worker". The grader says: partly answered — we have a service, not an owner. - Round 2. Refined query: "
refund-workerowner team on-call". It retrieves the service catalogue entry: owned by the Payments Reliability team, with a Slack channel. - Round 3 is not needed; the grader confirms both facts are present.
The answer cites both pages. A single-pass pipeline had returned only the deprecation page and replied, "The v1 API is deprecated; please migrate", which did not answer the question. The team caps the loop at three rounds, and the logs show that about 15% of questions use a second round and fewer than 3% use a third.
Follow-up questions to expect
- "How does the grader know what's missing?" — Ask it for structured output:
answered: boolplusmissing: str. Themissingfield becomes the next query. - "How is this different from RAG Fusion?" — RAG Fusion runs several queries in parallel, once. Iterative refinement runs queries in sequence, each informed by the last result — necessary for multi-hop questions.
- "How do you evaluate multi-hop retrieval?" — Label the full set of supporting documents per question and measure whether all of them were retrieved, not just one. Also track the number of rounds and latency per question.