Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG system indexes product documentation updated 20+ times a day. Users get outdated answers within hours. How do you keep the vector index fresh without re-embedding the entire corpus every time?
What you need to know
A typical documentation edit changes one paragraph. If that triggers re-embedding 300,000 chunks, you pay for 299,995 embeddings that produce identical vectors, and the job takes so long that answers stay stale anyway.
Chunk-level diffing
1from hashlib import sha25623def sync_document(doc_id: str, text: str, version: int):4 new = {f"{doc_id}:{i}": c for i, c in enumerate(chunk(text))}5 old = store.get_digests(prefix=f"{doc_id}:") # {chunk_id: digest}6 for cid, c in new.items():7 digest = sha256(normalise(c).encode()).hexdigest()8 if old.get(cid) != digest:9 store.upsert(cid, embed(c), digest=digest, doc_version=version)10 orphans = set(old) - set(new)11 store.delete(ids=list(orphans)) # removed sectionsThree parts matter:
- Stable ids built from the document id and position, so the same chunk has the same id next time.
- Normalisation before hashing (trim whitespace, unify line endings) so formatting noise doesn't count as a change.
- Orphan deletion. When a section is removed, its chunks must go too. Stale orphans are why users still see a discontinued price months later.
The pipeline
- Webhook — the docs CMS fires an event on every publish.
- Queue — SQS or Kafka holds the event; bursts and embedding rate limits are absorbed here, and failed jobs retry.
- Worker — fetches the document, chunks, diffs hashes, embeds only changed chunks in batches of about 100.
- Upsert and delete — writes to the vector store with a
doc_version. - Nightly reconciliation — re-hashes the whole corpus and repairs anything the webhooks missed.
Failure modes
| Failure | Symptom | Fix |
|---|---|---|
| Partial update | A document is half old, half new during the write | Write with doc_version; retrieval filters to the current version |
| Embedding rate limits | Worker errors during a bulk edit | Batch calls; the queue retries with backoff |
| Lost webhook | One document never updates | Nightly reconciliation job |
| Every id shifts after an insert near the top | A one-line edit re-embeds the whole document | Anchor chunk ids to headings, not just position |
The last row matters: with pure position ids, inserting a paragraph near the top shifts every later chunk's index, so all hashes "change". Using the section heading in the id (doc_id:installation:2) keeps most ids stable.
The metric
Index staleness: the p95 time from a source edit to the change being searchable. Set an SLO, say 5 minutes, and alert above it. Also count chunks in the index whose source no longer exists.
A real-life example
Scenario (illustrative numbers). A SaaS payments company has 4,200 documentation pages, about 180,000 chunks. A nightly full re-embed takes 3 hours and costs about $40 a night, and prices changed at 11 AM are still wrong in the bot until the next morning.
The team adds hashing, a webhook and a queue. A typical edit now re-embeds 3 to 6 chunks. Across 25 edits a day they embed about 120 chunks instead of 180,000. p95 edit-to-searchable drops from about 14 hours to 90 seconds. The first nightly reconciliation finds 2,300 orphaned chunks from pages deleted over the past year, including an old fee table that was still being quoted.
Follow-up questions to expect
- "What if the embedding model changes?" — Then every vector must change; that is a full migration with a parallel index, not an incremental update.
- "How do you handle deletes of whole documents?" — The CMS sends a delete event; the worker removes all chunks with that
doc_idprefix. - "Why keep the nightly full job at all?" — Webhooks get lost. The reconciliation is a cheap safety net that turns silent drift into a count you can alert on.