Course Content
Advanced RAG
3 sections · 38 lessons
How can you switch embedding models in production without downtime?
What you need to know
Why no in-place migration
An embedding model defines its own vector space. A query embedded with model B compared against documents embedded with model A gives meaningless similarities — often without any error. Even models from the same provider with the same dimension are not compatible unless the provider says so. So every document must be re-embedded, and the switch must be atomic from the user's point of view.
The migration
- Build alongside — create a new collection (or namespace, or table) for model B. Re-embed the corpus in a low-priority batch job. The live index is never touched.
- Dual-write — while the backfill runs, every document change goes to both indexes, so nothing is missed.
- Shadow-read — send live queries to both; serve A; log both result lists. Compare recall on your labelled set and overlap on real traffic.
- Canary — serve B to a small percentage of users or selected tenants, watching answer quality, feedback, latency and cost.
- Cut over — switch an alias or config flag, not a code deploy, so rollback takes seconds.
- Retire — keep A for a retention window, then delete it and stop dual-writing.
Comparing results in the shadow phase
1def overlap_at_k(old, new, k=5):2 return len(set(old[:k]) & set(new[:k])) / k34shadow_log = [ # (query language, old index top-5, new index top-5)5 ("en", ["a1", "a7", "a3", "a9", "a2"], ["a1", "a3", "a7", "a2", "a8"]),6 ("hi", ["b4", "b2", "b8", "b1", "b6"], ["b9", "b4", "b3", "b5", "b2"]),7 ("en", ["c2", "c5", "c1", "c3", "c4"], ["c2", "c5", "c1", "c4", "c7"]),8 ("ta", ["d3", "d1", "d6", "d2", "d8"], ["d7", "d9", "d3", "d4", "d5"]),9]10by_lang = {}11for lang, old, new in shadow_log:12 by_lang.setdefault(lang, []).append(overlap_at_k(old, new))13for lang, xs in by_lang.items():14 print(lang, round(sum(xs) / len(xs), 2)) # en 0.8, hi 0.4, ta 0.2Overlap does not say which model is better — it says where they disagree. English results barely change; Hindi and Tamil results change a lot. Those are the queries to label and inspect, because that is where the new model will either win or break things.
Details that bite
- Dimensions — a new dimension means a new schema. In pgvector a
vector(1024)column cannot hold 1,536-dimension vectors. - Distance metric and normalisation — use what the model was trained for (cosine or dot product). A wrong choice degrades quietly.
- Query prefixes — some models expect "query:" and "passage:" style prefixes or task instructions. Missing them lowers quality without errors.
- Thresholds — every similarity threshold (no-answer floor, semantic-cache threshold, dedup cutoff) must be re-tuned. A 0.82 that meant "relevant" in model A means nothing in model B.
- Derived data — semantic cache entries, stored retrieval results, and hypothetical-question vectors all belong to the old space. Rebuild or invalidate them.
- Cost and time — size the re-embedding job (tokens × price, or GPU hours) before committing.
A real-life example
An Indian telecom's support bot uses an English-centred embedding model. Hindi and Tamil queries often retrieve the wrong help article. The team wants to move to a stronger multilingual model with a different dimension.
- Week 1: a new collection is created; 12,000 articles and their generated questions are re-embedded overnight. A dual-write hook sends every article update to both collections.
- Weeks 2–3: shadow reads on all traffic. Overlap@5 is high for English and low for Hindi and Tamil. The team labels 400 Hindi, Tamil and Hinglish queries where the two disagree; the new model finds the right article far more often. But it also finds a regression: queries containing internal plan codes such as
PP-5G-399rank the right article lower, so they raise the weight of the BM25 leg for queries containing codes. - Week 4: canary to 5% of traffic in two circles, then 25%, then 100%. The semantic cache is rebuilt, and the no-answer threshold is re-tuned from 0.78 to 0.64 using the labelled set.
- Week 6: the old collection is deleted.
No customer sees downtime, and the rollback switch was never needed — but it was tested twice.
Follow-up questions to expect
- "Can you avoid re-embedding the whole corpus?" — Not safely, in general. Research on mapping one space to another exists, but for production the standard answer is full re-embedding.
- "What if the corpus is huge?" — Re-embed in priority order (most-retrieved documents first), throttle the job, and consider keeping the old index serving until coverage is complete.
- "How do you know the new model is really better?" — Your own labelled set, split by segment (language, query type), plus canary feedback. Public benchmark scores are a starting point, not proof for your domain.