Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your production RAG has 50 million embeddings from ada-002. A new model is 40% more accurate. You cannot re-index overnight — users are live. How do you migrate embeddings without downtime?
What you need to know
Why this cannot be done in place
Each embedding model has its own coordinate system. A query embedded with the new model and compared to chunks embedded with ada-002 gives meaningless similarity scores, without any error. Scores from two models are not on the same scale either, so you cannot merge results by score. The new model needs a complete, separate index. (OpenAI's text-embedding-3 models replaced ada-002 in 2024, and other providers and open models are common choices too.)
The plan
- Second collection — built with the new model; the old one keeps serving 100% of traffic.
- Status table — one row per chunk:
doc_id,chunk_hash,status,attempts. - Batch backfill — batch APIs are usually about half the price and latency does not matter. Estimate the token cost up front and get it approved.
- Dual-write — every create, update and delete goes to both indexes from the start.
- Shadow-read — a slice of live queries runs on both; compare against about 500 golden query-to-relevant-chunk pairs.
- Canary and roll back — 5%, 25%, 100%, watching quality, p95 latency and thumbs-down per arm; keep the old index for two weeks.
1CREATE TABLE reembed_status (2 chunk_id text PRIMARY KEY,3 chunk_hash text NOT NULL, -- skip work if the text has not changed4 status text NOT NULL DEFAULT 'pending', -- pending | done | failed5 attempts int NOT NULL DEFAULT 0,6 updated_at timestamptz NOT NULL DEFAULT now()7);89-- each worker claims a batch without clashing with other workers10SELECT chunk_id FROM reembed_status11WHERE status = 'pending' OR (status = 'failed' AND attempts < 5)12ORDER BY chunk_id LIMIT 200013FOR UPDATE SKIP LOCKED;FOR UPDATE SKIP LOCKED lets many workers pull batches in parallel without taking the same rows, and the chunk_hash column means a chunk edited during the backfill is re-embedded only once.
Use the moment well
| Decision | Why now |
|---|---|
| Fix chunking | You are paying to re-embed everything anyway |
| Choose the dimension | Some newer models support shorter vectors (Matryoshka-style truncation); measure the recall you lose against the memory you save |
| Re-check index settings | A different dimension changes memory and build time |
A real-life example
Scenario, numbers made up. An e-commerce marketplace holds 50M product-description chunks embedded with ada-002. A vendor says its new model is "40% more accurate". Re-embedding at their rate limit will take about four days.
The team creates the status table and turns on dual-write before starting. On day two, a provider outage fails 3M chunks; the workers retry them automatically the next morning with no restart. While re-embedding, they also fix chunking so size charts are no longer split from their products. Shadow reads on their 500-pair golden set show recall@10 up from 68% to 79% — a solid gain, though well short of the headline number. They canary over a week and drop the old collection after 14 days.
Follow-up questions to expect
- "How do you estimate the cost before starting?" — Count tokens on a sample of chunks, multiply to the full corpus, and price it at the batch rate; include a margin for retries.
- "What if the new model is worse on some query types?" — The shadow comparison shows it by query segment; you can hold the rollout, tune, or keep hybrid keyword search weighted higher for those segments.
- "Can you serve both indexes and merge results?" — Only with rank-based fusion such as RRF, never by raw scores, and it doubles query cost; it is a stop-gap, not the goal.