RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is Agentic RAG?


The corrective loop on a vague lounge questionretrieve,attempt 1grade:nothing relevantrewrite the queryretrieve,attempt 2grade: passanswerwith citationnullPlatinum loungevisits per year8 visits a year [1]After 3 failed attempts the graph ends with a clear not-found reply.
The agent's value is the decision after each retrieval; the attempt limit is what keeps that decision from looping forever.

What you need to know

Classic RAG makes one decision in advance: always retrieve k chunks from one index, then answer. Agentic RAG adds control flow, meaning the path through the system depends on the question and on intermediate results.

Classic RAG

  • One retrieval, fixed k, one index
  • Same steps for every question
  • 1 model call, predictable latency
  • Fails on multi-part or vague questions

Agentic RAG

  • Model picks sources, queries and number of steps
  • Path depends on the question
  • 2 to 10 model calls, variable latency
  • Handles comparisons, multi-hop, follow-ups

The four abilities

  • Routing: choose the right source for the question: product FAQ index, rate table in SQL, or web search.
  • Query rewriting: turn "and for Gold?" into "What is the annual fee for the Gold credit card?", or split a compound question into parts.
  • Grading: judge whether the retrieved chunks actually answer the question; if not, try again differently.
  • Multi-hop: use what one search found to write the next search.

Building it as a graph

Most teams build agentic RAG as a state graph (for example with LangGraph): nodes are steps, edges decide what runs next, and a shared state holds the question, query, documents and counters. A simple corrective loop:

Python
from typing import TypedDictfrom langgraph.graph import StateGraph, START, ENDclass State(TypedDict):    question: str    query: str    docs: list[str]    attempts: int    answer: strdef grade(s: State) -> str:     # a small LLM or reranker score in practice    if s["docs"]:        return "generate"    return "rewrite" if s["attempts"] < 3 else "give_up"g = StateGraph(State)for name, fn in [("retrieve", retrieve), ("rewrite", rewrite),                 ("generate", generate), ("give_up", give_up)]:    g.add_node(name, fn)g.add_edge(START, "retrieve")g.add_conditional_edges("retrieve", grade, ["generate", "rewrite", "give_up"])g.add_edge("rewrite", "retrieve")g.add_edge("generate", END)g.add_edge("give_up", END)app = g.compile()

retrieve increases attempts each time it runs. grade is the decision: good documents go to generate; empty or poor ones go to rewrite and back to retrieve; after 3 attempts the graph stops with a clear "not found". I ran this with stub nodes: the vague question "how many times can I use the airport lounge?" found nothing, was rewritten to "Platinum card lounge visits per year", and was answered on the second attempt.

The attempt limit is the most important line. Without it, a question with no answer loops until something times out.

When to use it

Use it for question types where your golden set shows single-pass retrieval failing: comparisons, questions across several documents, questions needing both documents and live data. Keep simple lookups on the fast path. A small classifier or rule at the start can decide which path each question takes.

A real-life example

A bank's FAQ bot handles 40,000 questions a day. Most are simple ("what is the IFSC code for the Andheri branch?"). But some ask things like "Which is cheaper for me, the Gold or Platinum card, if I spend Rs 40,000 a month?" That needs both cards' fee rules, both waiver thresholds and a calculation.

Single-pass retrieval returns chunks about one card and the answer compares wrongly. The team adds a router: questions classified as comparisons (about 6% of traffic) go to an agentic graph that retrieves each card's fee chunk separately, calls a small calculator tool, and then answers with citations. The other 94% stay on the one-call path. Average cost per question rises only slightly, and the comparison questions go from mostly wrong to mostly right on the 60 comparison cases in the golden set.

Follow-up questions to expect

  • "How is this different from a normal agent?" — It is an agent whose main tools are retrievers. The same risks apply: loops, cost, and untrusted tool output.
  • "How do you debug it?" — Tracing is essential; each run takes a different path. Log every decision (route chosen, grade, rewritten query) so you can see why it went where it went.
  • "Does it always improve quality?" — No. On simple questions it adds delay and more chances to go wrong. Measure it per question type.