Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

When should long-context models replace RAG, and how do you evaluate that decision?


Decide per question type, not onceLong context wins• Whole-dossier consistency checks• Small, stable, shared corpus• Facts spread across distant sections• Prototyping before building retrievalRetrieval wins• 400,000 documents, growing daily• Per-user access control on documents• Cost and latency at high traffic• Exact citations for every claim
Long context rarely deletes the retriever; it lets you retrieve more loosely and keep the retriever for size, freshness and permissions.

What you need to know

Where long context genuinely wins

  • Whole-document tasks. "Summarise this contract", "list every place this dossier contradicts itself", "trace this theme through the report". You can't know in advance which passages matter, and chunked retrieval misses connections across distant sections.
  • Small, stable, shared knowledge. A handbook or product catalogue that fits easily and is the same for everyone. Add prompt caching and you get simplicity with no retrieval misses.
  • Prototyping. Paste the documents in first; build retrieval when you understand the questions.

Where retrieval still wins

  • Size. A company's knowledge base is gigabytes. No window holds it, and that won't change soon.
  • Cost and latency. Prefill grows with input length. A 200k-token prompt per question costs vastly more than 4k tokens of retrieved context; caching only helps when the same prefix repeats often.
  • Access control. Retrieval filters by the user's permissions per query. One shared long context cannot serve users entitled to different documents.
  • Freshness. One changed document means re-sending, and re-caching, everything.
  • Attribution. Retrieval gives you chunk IDs for citations; with long context you rely on the model to cite precisely.
  • Accuracy at length. Needle-in-a-haystack tests (find one planted sentence) look excellent, but tasks that need several facts among many similar distractors degrade well before the advertised limit.

How to evaluate the decision

  1. Build one question set from real use, including single-fact lookups, multi-fact questions, whole-document questions, and questions whose answer is not in the corpus.
  2. Run three configurations: RAG (top 5), long context (everything that fits), and hybrid (retrieve top 50 into a large window).
  3. Score answer accuracy against references, faithfulness to sources, citation correctness, and refusal quality on unanswerable questions.
  4. Measure p50 and p95 latency and cost per query at expected traffic, with and without caching.
  5. Decide per question type, not globally; a router can send whole-document tasks to long context and lookups to RAG.

If you use an LLM judge, control its biases: it tends to prefer longer and more confident answers, and in side-by-side comparisons it is swayed by which answer comes first. Randomise order, blind the configuration names, and calibrate against human labels.

The usual outcome

Long context rarely deletes the retriever. It makes retrieval precision less critical: you can retrieve 30–50 chunks instead of 5, and sometimes drop the compression stage or loosen the reranker. The retriever stays, because it is what handles size, permissions and freshness.

A real-life example

A pharma company's regulatory team has two needs.

Dossier review. Before a submission, reviewers check a 250-page module for internal inconsistencies — a shelf life stated as 24 months in one section and 36 in another. RAG struggles: the two statements are in different chunks, and nothing in the question points to either. The team sends the whole module (about 150,000 tokens) to a long-context model with instructions to list every inconsistency with page references. On 20 past dossiers where reviewers had logged the real inconsistencies, long context finds most of them; RAG finds few. Each review costs a few dollars and takes minutes — trivial compared with reviewer time.

Regulatory search. For everyday questions across 400,000 documents with product-level access control, long context is not even possible. RAG stays, with hybrid retrieval and reranking.

For a middle case — questions about one product's full history (about 2 million tokens of documents) — they test RAG top-5, and a hybrid of RAG top-40 into a 200k window. The hybrid is more accurate on multi-fact questions, at about three times the cost per query. They route only "history and comparison" questions to the hybrid.

Follow-up questions to expect

  • "Won't bigger windows eventually make RAG obsolete?" — Not for large, changing, permissioned corpora. Cost, latency, freshness and access control are structural reasons, not temporary ones.
  • "How does prompt caching change the maths?" — It makes a repeated, shared prefix cheap and fast, which favours long context for small, stable corpora with steady traffic. It does nothing for per-user or rapidly changing content.
  • "What about 'lost in the middle'?" — Newer models handle position better, but long prompts full of similar passages still hurt. Put key material at the start or end, and test on your own data.