Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How does InstructRAG use reasoning steps to improve outputs?


What you need to know

The problem: noisy retrieval

Retrievers are noisy by design. A top-5 list usually contains one or two passages that merely look similar — an older policy, a document about a related product. A model asked only for the answer tends to anchor on the first plausible passage. InstructRAG makes "which evidence do I trust, and why?" an explicit step.

How the rationales are made

The clever part is that no human writes them:

  1. Collect training questions with gold answers, and run your retriever to get realistic, noisy documents for each.
  2. Synthesise — prompt an instruction-tuned model: "Here is the question, the documents and the correct answer. Explain how the answer is derived, and point out any documents that are irrelevant or wrong."
  3. Use them — either as a few in-context demonstrations (the training-free version, often called InstructRAG-ICL) or as supervised data to fine-tune the generator (InstructRAG-FT).
  4. Infer — at run time the model receives a question and documents only, and writes its own rationale, then the answer.

Giving the synthesiser the correct answer is what makes the rationales reliable: the model explains a known result instead of guessing one.

Why it helps

  • The model must look at every document and say what it is worth, instead of skimming the top one.
  • Contradictions become visible ("Doc 2 is the 2023 version; Doc 4 supersedes it").
  • The rationale is an audit trail. In regulated work, a reviewer can see why an answer was given.

Costs and limits

  • More output tokens, so more latency and cost per answer.
  • Fine-tuning needs training infrastructure and a model you can train.
  • The gain shrinks as retrieval gets cleaner. Test it on a deliberately noisy retrieval set, or you will conclude it does nothing.

Note a name clash: a separate 2025 paper also called "InstructRAG" is about task planning for agents. In interviews this question almost always means the rationale method above; if unsure, say which one you are describing.

A real-life example

A pharma company's regulatory search assistant often retrieves several versions of the same guideline — a draft, the final version and a later revision. For "What is the reporting timeline for serious unexpected adverse reactions in clinical trials?", the top 5 includes a draft with an old timeline.

The team builds 300 training examples from past questions answered by regulatory staff, retrieves documents for each, and has a model write rationales given the correct answers. A typical rationale:

Text
Doc 1 (final guideline, 2024) states the timeline directly and is current.Doc 3 is a 2019 draft with a different timeline; it is superseded by Doc 1.Doc 5 is about post-marketing reporting, not clinical trials; not relevant.Answer based on Doc 1.

They first try the training-free version with four such rationales as examples in the prompt. On a test set built with noisy retrieval, wrong-version answers fall clearly. Fine-tuning a smaller open model on all 300 gives similar accuracy at lower cost per query. Reviewers say the rationale is the feature they value most, because they can check it in seconds.

Follow-up questions to expect

  • "Isn't this just chain-of-thought?" — It is chain-of-thought focused on the evidence, and the key contribution is how the training rationales are generated: from the gold answer, so they are consistent and cheap to produce.
  • "Would a reranker make this unnecessary?" — A good reranker removes much of the noise, so the gain shrinks. But a reranker cannot tell that two relevant documents contradict each other; the rationale can.
  • "Do reasoning models need this?" — They already reason, but an explicit, short evidence rationale in the output is still valuable as an audit trail, and prompting for it costs little.