Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG system indexes product documentation updated 20+ times a day. Users get outdated answers within hours. How do you keep the vector index fresh without re-embedding the entire corpus every time?


One docs edit, a handful of embeddingsCMS webhookon publishQueue absorbsbursts and rate limitsRe-chunk,comparetext hashesUpsert 3 to 6changed chunksDelete orphansby doc id prefixA nightly job re-hashes everything to catch lost webhooks.
Deleting removed chunks is the step teams forget, and it is why bots keep quoting prices that no longer exist.

What you need to know

A typical documentation edit changes one paragraph. If that triggers re-embedding 300,000 chunks, you pay for 299,995 embeddings that produce identical vectors, and the job takes so long that answers stay stale anyway.

Chunk-level diffing

Python
from hashlib import sha256def sync_document(doc_id: str, text: str, version: int):    new = {f"{doc_id}:{i}": c for i, c in enumerate(chunk(text))}    old = store.get_digests(prefix=f"{doc_id}:")             # {chunk_id: digest}    for cid, c in new.items():        digest = sha256(normalise(c).encode()).hexdigest()        if old.get(cid) != digest:            store.upsert(cid, embed(c), digest=digest, doc_version=version)    orphans = set(old) - set(new)    store.delete(ids=list(orphans))                          # removed sections

Three parts matter:

  • Stable ids built from the document id and position, so the same chunk has the same id next time.
  • Normalisation before hashing (trim whitespace, unify line endings) so formatting noise doesn't count as a change.
  • Orphan deletion. When a section is removed, its chunks must go too. Stale orphans are why users still see a discontinued price months later.

The pipeline

  1. Webhook — the docs CMS fires an event on every publish.
  2. Queue — SQS or Kafka holds the event; bursts and embedding rate limits are absorbed here, and failed jobs retry.
  3. Worker — fetches the document, chunks, diffs hashes, embeds only changed chunks in batches of about 100.
  4. Upsert and delete — writes to the vector store with a doc_version.
  5. Nightly reconciliation — re-hashes the whole corpus and repairs anything the webhooks missed.

Failure modes

FailureSymptomFix
Partial updateA document is half old, half new during the writeWrite with doc_version; retrieval filters to the current version
Embedding rate limitsWorker errors during a bulk editBatch calls; the queue retries with backoff
Lost webhookOne document never updatesNightly reconciliation job
Every id shifts after an insert near the topA one-line edit re-embeds the whole documentAnchor chunk ids to headings, not just position

The last row matters: with pure position ids, inserting a paragraph near the top shifts every later chunk's index, so all hashes "change". Using the section heading in the id (doc_id:installation:2) keeps most ids stable.

The metric

Index staleness: the p95 time from a source edit to the change being searchable. Set an SLO, say 5 minutes, and alert above it. Also count chunks in the index whose source no longer exists.

A real-life example

Scenario (illustrative numbers). A SaaS payments company has 4,200 documentation pages, about 180,000 chunks. A nightly full re-embed takes 3 hours and costs about $40 a night, and prices changed at 11 AM are still wrong in the bot until the next morning.

The team adds hashing, a webhook and a queue. A typical edit now re-embeds 3 to 6 chunks. Across 25 edits a day they embed about 120 chunks instead of 180,000. p95 edit-to-searchable drops from about 14 hours to 90 seconds. The first nightly reconciliation finds 2,300 orphaned chunks from pages deleted over the past year, including an old fee table that was still being quoted.

Follow-up questions to expect

  • "What if the embedding model changes?" — Then every vector must change; that is a full migration with a parallel index, not an incremental update.
  • "How do you handle deletes of whole documents?" — The CMS sends a delete event; the worker removes all chunks with that doc_id prefix.
  • "Why keep the nightly full job at all?" — Webhooks get lost. The reconciliation is a cheap safety net that turns silent drift into a count you can alert on.