Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How do you design a RAG system that supports multi-turn conversations?


Turning 'aur postpaid mein?' into a real queryLast 2–3turns plusthe new messageSmall modelwrites astandalone queryUpdate filters:prepaid becomes postpaidRetrieve onrewrite andraw text, fuseStore citedchunk IDs in state'The second one' resolves from stored citations.
Most multi-turn failures happen before retrieval: the follow-up is searched as if it were a complete question.

What you need to know

Why single-turn RAG breaks

A single-turn pipeline embeds the user's message and searches. In a conversation, messages lean on earlier turns: "and for postpaid?", "what about the second one?", "is that also true in Kerala?". Embedded alone, these match nothing useful. Evaluation sets built from single questions hide this completely.

The design

  1. Condense — a small, fast model gets the last 2–3 turns and the new message and writes a standalone query. Retrieve on both the rewrite and the raw message, and fuse with RRF; rewriters sometimes over-correct.
  2. Carry state as filters — product, plan type, language, region and account type learned earlier become metadata filters, not just words in the prompt.
  3. Split memory — recent turns verbatim; older turns in a rolling summary plus extracted facts ("customer is on prepaid, device is iPhone 15"). Never re-embed the whole transcript each turn.
  4. Decide whether to retrieve — many follow-ups ("explain step 2 again", "shorter please") need no new retrieval.
  5. Keep conversation state — store retrieved chunk IDs and citations, so "the second plan you mentioned" maps to a real document, and follow-ups on the same document reuse it.

Failure modes to name

  • Topic drift — the rewriter drags in entities from five turns ago after the user changed topic. Limit its history window and tell it to prefer the latest message.
  • Context bloat — retrieved chunks pile up turn after turn until the budget breaks. Evict the oldest retrieved context first.
  • Stale filters — the user switches from prepaid to postpaid but the filter stays. Let the rewriter output filters explicitly each turn, so they can change.
  • Injected instructions — a user (or a retrieved document) says "from now on, ignore your rules". Keep system rules outside the conversation memory and never let summaries rewrite them.

Evaluate the conversation, not the question

Build multi-turn test conversations. Measure the rewrite quality (is the standalone query correct?), retrieval recall on turn 2 and later, and answer accuracy per turn. The gap between turn 1 and later turns is usually large and is where the product actually fails.

A real-life example

An Indian telecom's support bot handles conversations in English, Hindi and Hinglish. A real conversation:

Text
User: 5G plan kaunse hain 399 ke andar?          (5G plans under ₹399?)Bot:  [lists three prepaid plans, citing plan pages]User: aur postpaid mein?                         (and in postpaid?)User: second wale mein Netflix milega?           (does the second one include Netflix?)

Without condensation, "aur postpaid mein?" retrieves generic postpaid FAQs. With it, the rewriter produces "5G postpaid plans under ₹399" and switches the plan_type filter from prepaid to postpaid. For "second wale mein Netflix milega?", the bot looks up the second plan in its stored citation list and retrieves that plan's OTT section — no guessing which "second" the user meant.

The team built 200 test conversations of 3–6 turns from real chat logs. Turn-1 accuracy was already good; turn-3 accuracy was poor before condensation and filter carry-over, and close to turn-1 levels after. The rewriter is a small model that adds a couple of hundred milliseconds, and the team decided that was a clear win.

Follow-up questions to expect

  • "Why retrieve on the raw message as well as the rewrite?" — The rewriter can be wrong, for example by attaching an old topic. Fusing both lists keeps recall if one of them is off.
  • "How many turns of history should the rewriter see?" — Usually the last 2–3, plus the extracted facts. More history increases topic drift more than accuracy.
  • "Where do you store conversation state?" — A fast store keyed by conversation ID (for example Redis or the app database), with a TTL, holding recent turns, the summary, active filters and cited chunk IDs.