Course Content
RAG Systems
12 sections · 66 lessons
What happens if different embedding models are used for indexing and querying?
What you need to know
Each embedding model is trained separately. Dimension 17 in one model has nothing to do with dimension 17 in another. Comparing across models is like comparing GPS coordinates with postcodes: both are numbers, and the comparison is meaningless.
A real test: two 384-dimension models
Both all-MiniLM-L6-v2 and bge-small-en-v1.5 output 384 numbers, so a vector store accepts either without complaint. On a small HR handbook split into 45 chunks, with 10 labelled questions:
| Documents embedded with | Queries embedded with | Recall@1 | Average top score |
|---|---|---|---|
| MiniLM | MiniLM | 0.9 | 0.597 |
| BGE | BGE | 0.8 | 0.739 |
| BGE | MiniLM | 0.7 | 0.206 |
| MiniLM | BGE | 0.2 | 0.231 |
Two lessons from this toy run:
- It may not look random. One mixed direction still got 7 of 10 right, probably because these two models were trained on similar data. That is dangerous: a quick smoke test can pass while quality is broken for real traffic.
- The scores tell you. The best match's score dropped from about 0.6–0.74 to about 0.21–0.23. A sudden, across-the-board drop in top scores is a strong alarm that something changed in the embedding path.
Everything that must match
- The model name and version (a provider's "v2" is a new space).
- Normalisation (both normalised or both not).
- Prefixes and instructions (
query:on queries andpassage:on documents, if the model uses them). - Truncation and pre-processing such as lowercasing.
- Dimension truncation for Matryoshka models (both cut to the same size).
Enforcing it
1EMBED_MODEL = "BAAI/bge-small-en-v1.5"23store = Chroma(collection_name="hr_v7", embedding_function=emb,4 collection_metadata={"embed_model": EMBED_MODEL, "hnsw:space": "cosine"})56meta = store._collection.metadata or {}7assert meta.get("embed_model") == EMBED_MODEL, "query model differs from index model"Put the model name in the collection name or metadata, and fail loudly at start-up if they differ.
Changing models safely
- Build a new collection — re-embed every chunk with the new model into
docs_v8. - Evaluate — run the labelled set against both collections.
- Switch — point the alias or configuration at
docs_v8in one step. - Keep the old one briefly — so you can roll back, then delete it.
There is no way to convert old vectors into the new space reliably; plan for a full re-embed.
A real-life example
A bank's product-FAQ bot is upgraded: a developer changes the embedding model name in the query service to a newer version from the same provider. Both produce 1,536-dimension vectors. The ingestion job, in another repository, still uses the old model.
Nothing crashes. Over two days, the "helpful" rating on answers falls, and the share of "I could not find that" replies rises. The on-call engineer checks the retrieval logs and sees the average top similarity fell from about 0.55 to about 0.2 on the day of the deploy. Rolling back the query service fixes it in minutes. The team then adds the model name to collection metadata, a start-up assertion, and a dashboard alert on the daily average top score.
Follow-up questions to expect
- "Can you mix models in one collection?" — No. Use separate collections, query each with its own model, and fuse the ranked lists if you must.
- "What about a provider silently updating a model?" — Pin exact model versions where the provider offers them, and watch top-score drift as an early warning.
- "How long does a re-embed take?" — It scales with corpus size and throughput limits; plan it like a data migration with batching and retries.