Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your production RAG has 50 million embeddings from ada-002. A new model is 40% more accurate. You cannot re-index overnight — users are live. How do you migrate embeddings without downtime?


What you need to know

Why this cannot be done in place

Each embedding model has its own coordinate system. A query embedded with the new model and compared to chunks embedded with ada-002 gives meaningless similarity scores, without any error. Scores from two models are not on the same scale either, so you cannot merge results by score. The new model needs a complete, separate index. (OpenAI's text-embedding-3 models replaced ada-002 in 2024, and other providers and open models are common choices too.)

The plan

  1. Second collection — built with the new model; the old one keeps serving 100% of traffic.
  2. Status table — one row per chunk: doc_id, chunk_hash, status, attempts.
  3. Batch backfill — batch APIs are usually about half the price and latency does not matter. Estimate the token cost up front and get it approved.
  4. Dual-write — every create, update and delete goes to both indexes from the start.
  5. Shadow-read — a slice of live queries runs on both; compare against about 500 golden query-to-relevant-chunk pairs.
  6. Canary and roll back — 5%, 25%, 100%, watching quality, p95 latency and thumbs-down per arm; keep the old index for two weeks.
SQL
CREATE TABLE reembed_status (  chunk_id    text PRIMARY KEY,  chunk_hash  text NOT NULL,          -- skip work if the text has not changed  status      text NOT NULL DEFAULT 'pending',   -- pending | done | failed  attempts    int  NOT NULL DEFAULT 0,  updated_at  timestamptz NOT NULL DEFAULT now());-- each worker claims a batch without clashing with other workersSELECT chunk_id FROM reembed_statusWHERE status = 'pending' OR (status = 'failed' AND attempts < 5)ORDER BY chunk_id LIMIT 2000FOR UPDATE SKIP LOCKED;

FOR UPDATE SKIP LOCKED lets many workers pull batches in parallel without taking the same rows, and the chunk_hash column means a chunk edited during the backfill is re-embedded only once.

Use the moment well

DecisionWhy now
Fix chunkingYou are paying to re-embed everything anyway
Choose the dimensionSome newer models support shorter vectors (Matryoshka-style truncation); measure the recall you lose against the memory you save
Re-check index settingsA different dimension changes memory and build time

A real-life example

Scenario, numbers made up. An e-commerce marketplace holds 50M product-description chunks embedded with ada-002. A vendor says its new model is "40% more accurate". Re-embedding at their rate limit will take about four days.

The team creates the status table and turns on dual-write before starting. On day two, a provider outage fails 3M chunks; the workers retry them automatically the next morning with no restart. While re-embedding, they also fix chunking so size charts are no longer split from their products. Shadow reads on their 500-pair golden set show recall@10 up from 68% to 79% — a solid gain, though well short of the headline number. They canary over a week and drop the old collection after 14 days.

Follow-up questions to expect

  • "How do you estimate the cost before starting?" — Count tokens on a sample of chunks, multiply to the full corpus, and price it at the batch rate; include a margin for retries.
  • "What if the new model is worse on some query types?" — The shadow comparison shows it by query segment; you can hold the rollout, tune, or keep hybrid keyword search weighted higher for those segments.
  • "Can you serve both indexes and merge results?" — Only with rank-based fusion such as RRF, never by raw scores, and it doubles query cost; it is a stop-gap, not the goal.