Course Content
Advanced RAG
3 sections · 38 lessons
What is Corrective RAG (CRAG), and how does it improve retrieval accuracy?
What you need to know
The problem CRAG solves
A plain RAG pipeline has no idea whether retrieval succeeded. It always gets back some top-k list, because nearest-neighbour search always returns neighbours, even when none of them is relevant. The model then tries to use those passages. Irrelevant context does not get ignored: it pulls the answer towards whatever the passages say.
How the three actions are chosen
The evaluator scores every document. Two thresholds turn those scores into one action for the whole retrieval set:
1UPPER, LOWER = 0.7, 0.323def crag_action(scores: list[float]) -> str:4 best = max(scores)5 if best >= UPPER:6 return "correct" # at least one document is clearly relevant7 if best < LOWER:8 return "incorrect" # nothing is relevant: go to another source9 return "ambiguous" # unsure: refine internal and add external1011print(crag_action([0.91, 0.40, 0.12])) # correct12print(crag_action([0.22, 0.18, 0.05])) # incorrect13print(crag_action([0.55, 0.31, 0.20])) # ambiguousThe thresholds are not universal. You tune them on a labelled set of questions where you know whether the corpus contains the answer.
What each action does
- Correct → knowledge refinement. The paper calls it decompose-then-recompose: split each good document into small strips (a few sentences), score each strip, drop the weak ones, and join the rest. A relevant 800-token document often has only 150 useful tokens; the other 650 are distractors.
- Incorrect → external search. Rewrite the question into a search query and call another source. The paper used web search. In a company this is usually a second index, a broader search without filters, or an honest "I don't have that".
- Ambiguous → both. Keep the refined internal context and add the external results.
Where CRAG sits today
In 2026 most teams build CRAG as a graph of nodes — retrieve, grade, rewrite, search, generate — with a conditional edge after the grader (LangGraph and similar tools make this a few dozen lines). Batch the grading calls and use a cheap model, or the grader becomes the slowest step.
One security point: an external fallback brings untrusted text into the prompt. A web page can contain instructions like "ignore previous rules". Wrap external results in clear delimiters, tell the model they are data, and never let them trigger tools.
A real-life example
A bank builds a compliance assistant over its internal policy manuals. Officers ask questions like "What changed in the KYC re-verification rules for low-risk accounts?" The internal index is updated monthly, but regulator circulars arrive weekly.
In a test set of 300 questions, the team finds 40 where the internal index has no good document. Before CRAG, the assistant answered most of those confidently from older, similar-looking policies — the worst possible failure in compliance.
They add a grader (a small model returning a 0–1 score) and route "incorrect" cases to a second index built from the regulator's published circulars, not the open web. Anything still below the floor gets "I could not find a current source; please check with the compliance desk." Wrong confident answers on those 40 questions drop sharply, and every answer now cites either a policy or a circular.
Follow-up questions to expect
- "How is CRAG different from Self-RAG?" — Self-RAG trains the generator itself to emit reflection tokens that decide when to retrieve and whether the evidence supports the answer. CRAG is a plug-in evaluator; you do not retrain the generator.
- "Why not just use a reranker?" — A reranker only reorders; it still passes the top-k on. Using its score with a floor is actually the cheapest form of CRAG. If your corpus reliably contains answers, a reranker gives most of the benefit with less latency.
- "How do you evaluate it?" — Measure the grader like a classifier (precision and recall of "relevant"), and measure end-to-end how often the system answers wrongly versus refuses on questions the corpus cannot answer.