Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What role does planning play in PlanRAG?


What you need to know

The kind of question it is for

PlanRAG comes from a 2024 paper that studied LLMs as decision makers. The questions look like "Which product line should we discontinue?" or "Where should we open the next warehouse?" There is no document that answers these. The answer comes from a sequence of lookups — revenue by line, margin trend, contracts, inventory — combined with a judgement.

A similarity search over text is the wrong tool here. The data lives in tables or a graph, and the retrieval steps are queries (SQL, Cypher) rather than vector searches.

The three stages

  1. Plan — the model reads the question and the data schema, and writes the analysis steps: "1. Get alert counts per branch for the last two quarters. 2. Get the date of each branch's last audit. 3. Rank by alerts per 1,000 accounts, excluding branches audited in the last 6 months."
  2. Retrieve and answer — execute each step as a query, look at the result, and move on.
  3. Re-plan — if a result breaks an assumption (a table is empty, a column means something else, the data points elsewhere), write a new plan instead of forcing the old one.

The paper compared this with iterative RAG, which decides the next query one step at a time without an overall plan, and reported that the explicit plan did better on its decision-making benchmark. The intuition: without a plan, the model tends to stop after the first plausible lookup and miss factors it never thought to check.

What planning buys you

  • Decomposition. A vague business question becomes concrete, answerable queries.
  • Order. Step 3 can use the output of step 1.
  • Coverage. A written plan makes it obvious when a factor was skipped.
  • Auditability. The plan and each query's result form a trace a reviewer can check. For a recommendation someone must defend, that matters as much as the answer.

Costs and risks

  • Extra LLM calls before any data is fetched — add a few seconds.
  • A bad plan leads to a confident, well-formatted wrong answer. The re-plan step and a human review of high-stakes recommendations are the safeguards.
  • Generated SQL needs guardrails: read-only credentials, row limits, timeouts, and ideally a whitelist of views rather than raw tables.

For simple lookups such as "what is our refund window?", planning is pure overhead.

Where you see it now

Plan-then-execute is a standard agent pattern in 2026 (LangGraph and most agent frameworks have a template for it). PlanRAG is that pattern applied to retrieval over databases. Many "chat with your data" analytics assistants use this shape.

A real-life example

A bank's compliance head asks the assistant: "Which five branches should we prioritise for the next AML audit?"

The assistant's plan:

Text
1. Count suspicious-transaction alerts per branch, last 2 quarters.2. Normalise by number of active accounts per branch.3. Get each branch's last audit date and findings severity.4. Exclude branches audited in the last 6 months.5. Rank; explain the top 5 with their numbers.

Step 3 returns nothing for 40 branches. Instead of treating them as "never audited", the assistant re-plans: it checks the schema, finds that audits before 2024 are in an archive table, and adds a step to query it. Two branches that looked like top candidates drop out because they were audited recently.

The final answer lists five branches, each with the alert rate, last audit date and the queries used. The compliance head can check each number. Every SQL query ran through a read-only role limited to three reporting views.

Follow-up questions to expect

  • "How is PlanRAG different from iterative query refinement?" — Iterative refinement decides the next query based only on the last result. PlanRAG writes the full plan first and changes it only when the data contradicts it.
  • "How do you evaluate a planner?" — Check plan quality on labelled cases (did it include the needed factors?), query correctness (do they run and return the right rows?), and final-answer accuracy against an analyst's answer.
  • "What if the model writes a dangerous query?" — It can't, if the tool only has read-only access to approved views with row limits and timeouts. Never rely on the prompt for safety.