Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How does iterative query refinement improve answers?


Two hops to answer one questionQ: who ownsthe v1 ledgerAPI consumer?Round 1: findsrefund-workerGrader: servicefound, owner missingRound 2:refund-workerowner teamAnswer citesboth pagesCap rounds at 2–3 and stop when nothing new arrives.
The second query could not have been written before the first result came back — that is what makes a question multi-hop.

What you need to know

Why one pass is not enough

Some questions cannot be answered with one search, because the second search depends on the first result. "Who owns the service that still calls the deprecated v1 ledger API?" needs two hops: first find the service, then find its owner. No single chunk contains both facts.

The loop

  1. Retrieve with the current query.
  2. Grade — a small model or reranker threshold checks: is the question answered? What is missing?
  3. Refine — write a new query aimed at the missing piece.
  4. Accumulate — keep new documents, drop duplicates.
  5. Stop — when answered, after 2–3 rounds, or when a round adds nothing new. Then generate.

Kinds of refinement

  • Rewriting — fix vocabulary, expand acronyms, turn a follow-up ("and for enterprise?") into a standalone question.
  • Step-back prompting — ask a broader question first ("what rules govern data retention for logs?") to fetch the principle, then combine it with the specific query. Useful when the literal question is too narrow to match anything.
  • Self-query — have the model turn "incidents in the payments repo since March" into a filter repo = "payments" AND date >= "2026-03-01" plus a semantic query, so the index does the filtering.
  • Decomposition — split a compound question into sub-questions and retrieve for each.
  • Multi-hop — use an entity found in round 1 (a service name, a drug, a counterparty) as the query for round 2. Methods like IRCoT interleave these retrieval steps with chain-of-thought reasoning.

Agentic RAG

When an LLM runs this loop with retrieval as a tool, deciding itself what to search next and when to stop, it is called agentic RAG. Modern "deep research" features are the long version of this loop, running dozens of searches. For a product feature you want a much tighter version: a few rounds, a time budget, and clear logs of every query the agent issued.

Termination and cost

Every round adds a grading call, a rewrite call and a search — often 1–2 seconds. Without limits, a question the corpus cannot answer loops forever, reformulating each time. So set a maximum number of rounds, a total time budget, and a "no new documents" stop rule. When the loop ends without an answer, say so honestly.

A real-life example

A software company's engineering-wiki assistant gets this question from a new engineer:

Text
Who do I talk to about the service that still uses the v1 ledger API?
  • Round 1. Query: "service using v1 ledger API". It retrieves the ledger team's deprecation page, which says "remaining v1 consumers: refund-worker". The grader says: partly answered — we have a service, not an owner.
  • Round 2. Refined query: "refund-worker owner team on-call". It retrieves the service catalogue entry: owned by the Payments Reliability team, with a Slack channel.
  • Round 3 is not needed; the grader confirms both facts are present.

The answer cites both pages. A single-pass pipeline had returned only the deprecation page and replied, "The v1 API is deprecated; please migrate", which did not answer the question. The team caps the loop at three rounds, and the logs show that about 15% of questions use a second round and fewer than 3% use a third.

Follow-up questions to expect

  • "How does the grader know what's missing?" — Ask it for structured output: answered: bool plus missing: str. The missing field becomes the next query.
  • "How is this different from RAG Fusion?" — RAG Fusion runs several queries in parallel, once. Iterative refinement runs queries in sequence, each informed by the last result — necessary for multi-hop questions.
  • "How do you evaluate multi-hop retrieval?" — Label the full set of supporting documents per question and measure whether all of them were retrieved, not just one. Also track the number of rounds and latency per question.