Course Content
RAG Systems
12 sections · 66 lessons
What is self-reflection or self-correction in RAG systems?
What you need to know
There are two places to check: before generation (are these documents good?) and after generation (is this answer supported?).
Corrective RAG (CRAG)
From a 2024 paper by Yan and others. A retrieval evaluator scores each retrieved document for relevance, and the system takes one of three actions:
- Correct (good documents): refine them, keeping the useful sentences, and generate.
- Incorrect (poor documents): discard them and use another source; the paper used web search.
- Ambiguous: combine both.
The paper's evaluator was a small fine-tuned model, not a large LLM, which keeps it cheap.
Self-RAG
From a 2023 paper by Asai and others. The model is trained to output special reflection tokens that say whether retrieval is needed, whether a passage is relevant, whether each part of the answer is supported, and how useful the answer is. It uses these to decide and to pick among candidate answers.
This is a correction to a common simplification. Most production "Self-RAG" systems do not use the trained model. They copy the idea with prompted grader calls in a LangGraph flow: grade documents, generate, check "is this supported by the documents?", check "does this answer the question?", and loop once if not.
The flow
- Retrieve — first attempt, normal query.
- Grade documents — keep relevant ones; if none pass, rewrite the query and retrieve again, or switch source.
- Generate — from the kept documents.
- Check grounding — is each claim supported? If not, regenerate once with stricter instructions.
- Check usefulness — does it answer the question? If not, one more retrieval with a rewritten query.
- Stop — after the retry limit, return the best supported answer or "not found".
Cost arithmetic
Suppose the answer call costs 1 unit and a small grader call costs 0.1 units. Grading 5 documents one call each adds 0.5; the grounding and usefulness checks add 0.2. A clean pass costs about 1.7 units. If 20% of questions need one retry (another retrieval, grading and answer, about 1.7 more), the average is about 1.7 + 0.2 × 1.7 ≈ 2 units, roughly double plain RAG. Grading all documents in one batched call, or using reranker scores as the grade, cuts this a lot.
A real-life example
A hospital's guideline search often gets questions phrased the way nurses talk: "can we give the blood thinner before the scan?" First-pass retrieval sometimes returns general anticoagulation text that does not mention imaging.
The team adds a document grader (a small model asked: "Does this passage answer the question? yes or no") and a query rewriter. When no passage passes, the rewriter produces "anticoagulant administration before CT imaging guideline", which retrieves the right protocol. After two failed attempts, the tool replies "No guideline found for this situation; contact the on-call pharmacist." The grader's logs become the most useful data the team has: the most frequently failed topics show which guidelines are missing or badly chunked.
Follow-up questions to expect
- "What if the grader is wrong?" — It will be sometimes. Check its decisions against human labels, and treat a reject as "try another way", not as proof there is no answer.
- "How many retries?" — One or two. After that, quality rarely improves and cost keeps rising.
- "Can a reranker replace the LLM grader?" — Often yes. A reranker score with a tuned threshold is cheaper and more consistent than an LLM yes/no.