Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

When should retrieved content be summarized before being sent to the model?


What you need to know

Three ways to shrink retrieved context

Truncate or rerank

  • Drop whole low-ranked passages
  • No rewriting at all
  • Loses coverage, keeps accuracy

Extractive compression

  • Keep only query-relevant sentences, verbatim
  • Citations still point at real text
  • Moderate savings, low risk

Abstractive summary

  • An LLM rewrites content in fewer words
  • Biggest savings; enables synthesis
  • Can drop or distort qualifiers

Try them in that order. Summarisation is the last step, not the first.

When summarising is the right call

  • Synthesis over many documents. "What are customers complaining about this week?" has no single source. The answer is a summary. The standard pattern is map-reduce: summarise each document (or batch) for the question in parallel (map), then combine the partial summaries (reduce).
  • One huge document came back and only its relevance to this question matters.
  • Conversation memory — older turns are condensed into a running summary so recent turns and fresh retrieval fit.
  • Cost tiering — a cheap model condenses material before an expensive reasoning model reads it, so the premium rate applies to fewer tokens.

When not to summarise

  • Exact-wording domains — contract terms, dosage, regulations, prices. "Refundable within 30 days except for sale items" easily becomes "refundable within 30 days".
  • The reranked top-5 already fits — summarising adds an LLM round trip to the critical path for no gain.
  • You need precise citations — a summary's sentences don't map back to one source line.

If you summarise, do it safely

  • Summarise per document, in parallel, not over the concatenation — mixing sources is how attribution gets lost.
  • Keep the source ID attached to each summary, so the final answer can cite it.
  • Make it query-focused: "Summarise what this document says about X", not a generic summary.
  • Tell the summariser to copy numbers, dates, names and exceptions exactly.
  • Measure answer accuracy and faithfulness with and without the summary step. Token savings alone are not success.

A real-life example

An Indian telecom's operations team asks its support assistant every Monday: "What were the top reasons for complaints in the Tamil Nadu circle last week?" There are about 3,000 relevant tickets, in Tamil, English and Tanglish.

Retrieving the "top 10 most similar tickets" is useless — the answer is a distribution, not a document. So the pipeline runs map-reduce:

  1. Map: a small, cheap model summarises each ticket into one English line with a category and the ticket ID: "T-88213 — Billing: charged for an OTT add-on not requested."
  2. Reduce: batches of 200 lines are grouped into themes with counts and example IDs, then the batch results are merged.
  3. The final answer lists the top five themes with counts and links to example tickets, so a manager can check the source.

The same company's customer-facing bot takes the opposite approach for plan terms. When a customer asks about the fair-usage limit on a plan, the bot sends the exact plan-terms sentences (extractive), because a summary once turned "1.5 GB/day, then 64 kbps" into "1.5 GB/day data" and the bot wrongly told a customer their data would stop completely.

Follow-up questions to expect

  • "How do you check a summary didn't lose something important?" — Compare facts: extract numbers, dates and named entities from source and summary and flag missing ones; and run a faithfulness check with an LLM judge on a sample.
  • "Map-reduce or one long-context call?" — If everything fits and the model handles it well, one call is simpler. For thousands of documents, map-reduce scales, runs in parallel and keeps per-document traceability.
  • "Who should do the summarising?" — Usually a smaller, cheaper model for the map step; a stronger model for the final reduce and answer.