Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is Corrective RAG (CRAG), and how does it improve retrieval accuracy?


One grader score, three different actionsbest score0.7 or morerefine into stripsnobest 0.3 to 0.7refine into stripsyesbest under 0.3discard allyes, rewritten queryWhenInternal documentsExternal searchCorrectAmbiguousIncorrectThresholds are tuned on questions where you know the corpus lacks the answer.
Plain RAG always takes the top row; CRAG's value is noticing when it is really in the bottom one.

What you need to know

The problem CRAG solves

A plain RAG pipeline has no idea whether retrieval succeeded. It always gets back some top-k list, because nearest-neighbour search always returns neighbours, even when none of them is relevant. The model then tries to use those passages. Irrelevant context does not get ignored: it pulls the answer towards whatever the passages say.

How the three actions are chosen

The evaluator scores every document. Two thresholds turn those scores into one action for the whole retrieval set:

Python
UPPER, LOWER = 0.7, 0.3def crag_action(scores: list[float]) -> str:    best = max(scores)    if best >= UPPER:        return "correct"      # at least one document is clearly relevant    if best < LOWER:        return "incorrect"    # nothing is relevant: go to another source    return "ambiguous"        # unsure: refine internal and add externalprint(crag_action([0.91, 0.40, 0.12]))   # correctprint(crag_action([0.22, 0.18, 0.05]))   # incorrectprint(crag_action([0.55, 0.31, 0.20]))   # ambiguous

The thresholds are not universal. You tune them on a labelled set of questions where you know whether the corpus contains the answer.

What each action does

  • Correct → knowledge refinement. The paper calls it decompose-then-recompose: split each good document into small strips (a few sentences), score each strip, drop the weak ones, and join the rest. A relevant 800-token document often has only 150 useful tokens; the other 650 are distractors.
  • Incorrect → external search. Rewrite the question into a search query and call another source. The paper used web search. In a company this is usually a second index, a broader search without filters, or an honest "I don't have that".
  • Ambiguous → both. Keep the refined internal context and add the external results.

Where CRAG sits today

In 2026 most teams build CRAG as a graph of nodes — retrieve, grade, rewrite, search, generate — with a conditional edge after the grader (LangGraph and similar tools make this a few dozen lines). Batch the grading calls and use a cheap model, or the grader becomes the slowest step.

One security point: an external fallback brings untrusted text into the prompt. A web page can contain instructions like "ignore previous rules". Wrap external results in clear delimiters, tell the model they are data, and never let them trigger tools.

A real-life example

A bank builds a compliance assistant over its internal policy manuals. Officers ask questions like "What changed in the KYC re-verification rules for low-risk accounts?" The internal index is updated monthly, but regulator circulars arrive weekly.

In a test set of 300 questions, the team finds 40 where the internal index has no good document. Before CRAG, the assistant answered most of those confidently from older, similar-looking policies — the worst possible failure in compliance.

They add a grader (a small model returning a 0–1 score) and route "incorrect" cases to a second index built from the regulator's published circulars, not the open web. Anything still below the floor gets "I could not find a current source; please check with the compliance desk." Wrong confident answers on those 40 questions drop sharply, and every answer now cites either a policy or a circular.

Follow-up questions to expect

  • "How is CRAG different from Self-RAG?" — Self-RAG trains the generator itself to emit reflection tokens that decide when to retrieve and whether the evidence supports the answer. CRAG is a plug-in evaluator; you do not retrain the generator.
  • "Why not just use a reranker?" — A reranker only reorders; it still passes the top-k on. Using its score with a floor is actually the cheapest form of CRAG. If your corpus reliably contains answers, a reranker gives most of the benefit with less latency.
  • "How do you evaluate it?" — Measure the grader like a classifier (precision and recall of "relevant"), and measure end-to-end how often the system answers wrongly versus refuses on questions the corpus cannot answer.