Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

You've embedded 50M documents with ada-002. A new model is 2× better but incompatible with your existing vectors. How do you migrate embeddings without taking the system offline?


What you need to know

Why you cannot mix vectors

An embedding model turns text into a list of numbers. The meaning of each position is private to that model. A vector from text-embedding-ada-002 and one from a newer model (OpenAI's text-embedding-3 family replaced ada-002 in 2024) may even have different lengths, and when they do match, the numbers mean different things. Comparing them gives neighbours that look plausible and are meaningless, with no error. So the new model needs its own index, and every query must be embedded with the same model as the index it searches.

The migration plan

  1. Create a separate collection — name it with the model and dimension, e.g. docs__te3large_1024, so vectors cannot be mixed by mistake.
  2. Dual-write from day one — every create, update and delete goes to both indexes before the backfill starts.
  3. Backfill in batches — a queue of document IDs, workers that embed and upsert, and a checkpoint so a crash resumes instead of restarting. Use the provider's batch API, which is cheaper, because latency does not matter here.
  4. Shadow-read — send a sample of live queries to both indexes, log both result lists, and score them on your golden set.
  5. Ramp behind a flag — 1%, 10%, 50%, 100%, watching click-through, thumbs-down and escalations. Rollback is flipping the flag.
  6. Retire the old index — keep it for a couple of weeks, then delete it.

Dual-write is the step people forget. Without it, the backfill finishes with an index that is days out of date.

Python
INDEXES = {"old": ("docs__ada002_1536", embed_ada), "new": ("docs__te3large_1024", embed_new)}def upsert(doc):                               # called on every create or update    for name, (collection, embed) in INDEXES.items():        vdb.upsert(collection, id=doc.id, vector=embed(doc.text), payload=doc.meta)def search(query, user_id, k=10):    use_new = flags.percent("embeddings_v2") > hash_bucket(user_id)   # sticky per user    collection, embed = INDEXES["new" if use_new else "old"]    return vdb.search(collection, vector=embed(query), limit=k)       # same model both sides

Bucketing by user keeps each person on one index during the ramp, so their results do not flip between requests.

Planning the backfill

Throughput is set by your rate limit, not your code. At 800 documents per second, 50M documents take about 17 hours; at 200 per second, nearly three days. Plan for retries with backoff, and budget memory: two full indexes live side by side until cut-over.

ChoiceBenefitCost
Separate collectionClean rollback, no mixingDouble storage for a while
Shadow-read before rampProves the gain on your dataExtra query cost on the sample
Smaller output dimension (if the new model supports it)Less memory for the new indexSmall quality loss; measure it

A real-life example

Scenario, numbers made up. A legal-research platform holds 50M case-law passages embedded with ada-002. A newer model scores better on public benchmarks, and the team wants to switch without downtime.

They create a new collection and turn on dual-write on Monday. The backfill runs through the provider's batch API at about 800 passages per second and finishes in under a day. Shadow reads on 5% of traffic for a week show recall@10 on their 600-question golden set rising from 71% to 83% — less than the "2×" headline, but real. They ramp over ten days with no rise in thumbs-down, delete the old collection two weeks later, and memory drops because they chose a 1,024-dimension output.

Follow-up questions to expect

  • "Could you train a mapping from old vectors to new ones instead?" — A learned linear map can work as a stop-gap, but it loses quality and still needs testing. For a lasting migration, re-embed.
  • "What about deletes during the backfill?" — They go through dual-write too. The backfill must not re-insert a document deleted after its ID was queued, so check a tombstone or version before upserting.
  • "How do you know the new model is really better?" — Shadow-read on real queries and score both on your own golden set. Public benchmark gains often shrink on a specific domain.