Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How can you switch embedding models in production without downtime?


A parallel index and a reversible switchBuild new index,re-embed in backgroundDual-writechanges to bothShadow-read:serve old,log bothCanary 5, 25,then 100 percentFlip alias;re-tuneevery thresholdOverlap was high for English, low for Hindi and Tamil.
Shadow reads show where the two models disagree, which is exactly where the new one will either win or break things.

What you need to know

Why no in-place migration

An embedding model defines its own vector space. A query embedded with model B compared against documents embedded with model A gives meaningless similarities — often without any error. Even models from the same provider with the same dimension are not compatible unless the provider says so. So every document must be re-embedded, and the switch must be atomic from the user's point of view.

The migration

  1. Build alongside — create a new collection (or namespace, or table) for model B. Re-embed the corpus in a low-priority batch job. The live index is never touched.
  2. Dual-write — while the backfill runs, every document change goes to both indexes, so nothing is missed.
  3. Shadow-read — send live queries to both; serve A; log both result lists. Compare recall on your labelled set and overlap on real traffic.
  4. Canary — serve B to a small percentage of users or selected tenants, watching answer quality, feedback, latency and cost.
  5. Cut over — switch an alias or config flag, not a code deploy, so rollback takes seconds.
  6. Retire — keep A for a retention window, then delete it and stop dual-writing.

Comparing results in the shadow phase

Python
def overlap_at_k(old, new, k=5):    return len(set(old[:k]) & set(new[:k])) / kshadow_log = [   # (query language, old index top-5, new index top-5)    ("en", ["a1", "a7", "a3", "a9", "a2"], ["a1", "a3", "a7", "a2", "a8"]),    ("hi", ["b4", "b2", "b8", "b1", "b6"], ["b9", "b4", "b3", "b5", "b2"]),    ("en", ["c2", "c5", "c1", "c3", "c4"], ["c2", "c5", "c1", "c4", "c7"]),    ("ta", ["d3", "d1", "d6", "d2", "d8"], ["d7", "d9", "d3", "d4", "d5"]),]by_lang = {}for lang, old, new in shadow_log:    by_lang.setdefault(lang, []).append(overlap_at_k(old, new))for lang, xs in by_lang.items():    print(lang, round(sum(xs) / len(xs), 2))    # en 0.8, hi 0.4, ta 0.2

Overlap does not say which model is better — it says where they disagree. English results barely change; Hindi and Tamil results change a lot. Those are the queries to label and inspect, because that is where the new model will either win or break things.

Details that bite

  • Dimensions — a new dimension means a new schema. In pgvector a vector(1024) column cannot hold 1,536-dimension vectors.
  • Distance metric and normalisation — use what the model was trained for (cosine or dot product). A wrong choice degrades quietly.
  • Query prefixes — some models expect "query:" and "passage:" style prefixes or task instructions. Missing them lowers quality without errors.
  • Thresholds — every similarity threshold (no-answer floor, semantic-cache threshold, dedup cutoff) must be re-tuned. A 0.82 that meant "relevant" in model A means nothing in model B.
  • Derived data — semantic cache entries, stored retrieval results, and hypothetical-question vectors all belong to the old space. Rebuild or invalidate them.
  • Cost and time — size the re-embedding job (tokens × price, or GPU hours) before committing.

A real-life example

An Indian telecom's support bot uses an English-centred embedding model. Hindi and Tamil queries often retrieve the wrong help article. The team wants to move to a stronger multilingual model with a different dimension.

  • Week 1: a new collection is created; 12,000 articles and their generated questions are re-embedded overnight. A dual-write hook sends every article update to both collections.
  • Weeks 2–3: shadow reads on all traffic. Overlap@5 is high for English and low for Hindi and Tamil. The team labels 400 Hindi, Tamil and Hinglish queries where the two disagree; the new model finds the right article far more often. But it also finds a regression: queries containing internal plan codes such as PP-5G-399 rank the right article lower, so they raise the weight of the BM25 leg for queries containing codes.
  • Week 4: canary to 5% of traffic in two circles, then 25%, then 100%. The semantic cache is rebuilt, and the no-answer threshold is re-tuned from 0.78 to 0.64 using the labelled set.
  • Week 6: the old collection is deleted.

No customer sees downtime, and the rollback switch was never needed — but it was tested twice.

Follow-up questions to expect

  • "Can you avoid re-embedding the whole corpus?" — Not safely, in general. Research on mapping one space to another exists, but for production the standard answer is full re-embedding.
  • "What if the corpus is huge?" — Re-embed in priority order (most-retrieved documents first), throttle the job, and consider keeping the old index serving until coverage is complete.
  • "How do you know the new model is really better?" — Your own labelled set, split by segment (language, query type), plus canary feedback. Public benchmark scores are a starting point, not proof for your domain.