Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

What is federated RAG, and how does it query multiple sources?


Query sources where they live, as the real userRoute: whichsources can answer?Fan out with theuser's own tokenNormalise toone result shapeRRF, thendedupeand rerankAnswer withper-sourcecitationsA slow source times out; the answer says it was unavailable.
Scores from different sources cannot be compared, so rank fusion comes first and the reranker supplies the one real score.

What you need to know

Why not one big index

Copying everything into a single vector store is simpler, so federate only when you must. Common reasons:

  • Data residency or regulation — some data may not leave a region or a system.
  • Ownership — the legal team will not export its document system; the partner only offers an API.
  • Live data — stock levels or ticket status change every minute; a copy is stale on arrival.
  • Permissions — the source already enforces complex access rules that you do not want to re-implement.

The flow

  1. Route — decide which sources can answer this question (rules, a classifier, or an LLM planner).
  2. Fan out — query the chosen sources in parallel, each with a timeout, passing the caller's identity.
  3. Normalise — convert every result to one shape: text, source, URL or ID, timestamp, access label.
  4. Fuse and dedupe — merge the lists with RRF, drop near-duplicates.
  5. Rerank — a cross-encoder scores the merged set against the question on one common scale.
  6. Generate — answer with citations that name the source.

In 2026 the sources are often exposed as tools, frequently through MCP servers, so an agent can call search_policies, query_cases or search_wiki the same way. The engineering problems below do not change.

The hard parts

  • Score incomparability. BM25 scores, cosine similarities and a search API's opaque "relevance" are on different scales. That is why rank-based fusion (RRF) comes first and a reranker then gives one real score.
  • Tail latency. The answer waits for the slowest source. Give each source a timeout (say 800 ms) and answer from partial results, telling the user if a source was unavailable.
  • Duplicates. The same circular may exist in the policy store, the wiki and the email archive. Dedupe by hash or high similarity before reranking, or one fact fills the context three times.
  • Permissions. Pass the user's identity to every source (for example an on-behalf-of token) so each source returns only what that user may see. Never query with a super-user account and filter afterwards — one bug and you have leaked data.
  • Untrusted content. External sources can contain prompt-injection text. Mark retrieved text as data in the prompt, and do not let it trigger actions.

A real-life example

A bank's compliance assistant must answer questions such as "Has any customer flagged under the new PEP rules had a transaction above ₹10 lakh this quarter?" The knowledge is spread across:

SourceWhere it livesHow it is queried
Internal policy manualsVector store (hybrid search)Retrieval
Regulator circularsSeparate curated indexRetrieval
Case notesCase-management system with its own permissionsSearch API with the officer's token
TransactionsData warehouseRead-only SQL through a tool

A router sends the "PEP rules" part to the two document indexes and the transaction part to SQL. Case notes are queried with the officer's own token, so an officer from the retail team never sees notes from the corporate-banking investigations team. The case-management search is slow at month-end, so it gets a one-second timeout; when it times out, the answer says "case notes unavailable" rather than silently leaving them out.

The design review's key rule: the assistant has no account of its own that can read everything. Every source sees the real user.

Follow-up questions to expect

  • "Why not normalise scores to 0–1 and average them?" — Min-max normalisation depends on each result list's own spread and is unstable when a source returns few results. RRF plus a reranker is more robust.
  • "How do you decide which sources to query?" — Start with rules or a classifier trained on labelled questions; fan out to more sources when confidence is low, because an extra source costs little and a missed one costs the answer.
  • "How do you cite across sources?" — Keep a source ID and link on every normalised result and require the model to cite by ID. Validate that each cited ID was actually in the context.