Course Content
Advanced RAG
3 sections · 38 lessons
What is Modular RAG, and how does it differ from basic/advanced RAG?
What you need to know
The terms come from a widely cited 2023–2024 survey of RAG (Gao et al.), which grouped systems into these three "paradigms". Interviewers use them as shorthand, so know what each adds.
Naive RAG
chunk → embed → retrieve top-k → put in prompt → generateOne retrieval, one generation, no feedback. Its known failures: poor recall on paraphrased questions, irrelevant chunks in the prompt, and no way to notice that retrieval failed.
Advanced RAG
Same straight line, with extra stages bolted on:
- Pre-retrieval — query rewriting, multi-query, HyDE, metadata filter extraction.
- Retrieval — hybrid BM25 plus dense search fused with RRF.
- Post-retrieval — cross-encoder reranking, deduplication, context compression.
Still one pass. Every query takes every stage, whether it needs it or not.
Modular RAG
Concretely, this is what you build with a graph framework such as LangGraph: nodes for route, retrieve, grade, rewrite, generate, check_answer, and conditional edges between them. "If the grader says the context is weak, go back to rewrite — at most twice." Corrective RAG, Self-RAG-style checks and agentic RAG are all patterns built from these modules.
Advanced RAG
- A chain: every query runs every stage
- Predictable latency and cost
- One log line tells the story
- Improve it by tuning stages
Modular RAG
- A graph: path depends on the query
- Latency varies from 1 to many calls
- Needs per-node tracing
- Improve it by swapping or rerouting modules
Why the modularity matters in practice
- Per-query adaptivity. "What's the VPN URL?" skips rewriting and reranking; "Compare our two deployment guides" gets multi-query and a second retrieval round.
- Swap one piece. You can A/B a new reranker or embedding model behind the same interface without touching the rest.
- Evaluate each module. Retrieval recall, grader accuracy and answer faithfulness each get their own number.
The costs are real: more moving parts, loops that need hard limits, and p95 latency that can be several times the median. Without per-node traces (inputs, outputs, timings for every step) a bad answer is almost impossible to diagnose.
A real-life example
A software company has an engineering-wiki assistant over 30,000 pages: runbooks, design docs and API references. Version 1 is naive RAG. Engineers complain that it answers "How do I rotate the Kafka certificates?" with the 2023 runbook instead of the 2025 one.
Version 2 is advanced RAG: hybrid search, a reranker, and a last_updated boost. Accuracy improves, but every query now costs about 1.8 seconds, including "where is the on-call calendar?", which needs no reranking.
Version 3 is modular. A router sends navigation questions ("where is…") straight to a title search that returns a link in 200 ms. How-to questions take the full hybrid-plus-rerank path. When the grader finds that no chunk mentions the named service, the graph rewrites the query once, using the service's aliases from a lookup table, and retrieves again. Median latency falls, and the team can now test a new reranker on the how-to path alone.
Follow-up questions to expect
- "Is agentic RAG the same as modular RAG?" — Agentic RAG is a modular RAG where an LLM chooses the next step (which tool, whether to search again). Modular RAG can also be routed by fixed rules.
- "How do you stop loops from running away?" — A hard cap on iterations (2–3), a total time budget, and a rule to stop when a round returns no new documents.
- "When would you not go modular?" — Small, uniform workloads such as an FAQ bot. A well-tuned advanced RAG chain is cheaper to run and far easier to debug.