Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Your RAG system indexes product documentation updated 20+ times a day. Users get outdated answers within hours. How do you keep the vector index fresh without re-embedding the entire corpus every time?
What you need to know
This version of the question is really about targets and trade-offs: how fresh is fresh enough, what it costs, and how to protect answers while the index catches up.
The target decides the architecture
| Target | Design | Cost and complexity |
|---|---|---|
| Under a minute | Event-driven: webhook, stream (Kafka), always-on workers | Highest; needs on-call for the pipeline |
| About 5 to 10 minutes | Micro-batch: a job every few minutes processes changed documents | Moderate; simple to operate |
| Daily | Nightly incremental job | Cheapest; fine for slow-changing content |
For product docs edited 20+ times a day, a 5-minute target is usually enough, and a micro-batch or a small queue-based worker meets it.
Incremental indexing, briefly
Each chunk has a stable id (doc_id plus heading or position) and a stored hash of its text. An update re-chunks the document, compares hashes, upserts only changed chunks and deletes orphans. Twenty edits a day become a few hundred embeddings, not a few million. Use content-anchored splitting, on headings, so a one-word change near the top does not shift every chunk id and re-embed the whole document.
Protecting answers during change
- Recency boost at rank time. Store
updated_aton every chunk. When an old and a new chunk both match, add a small boost to the newer one:
1def rerank_with_recency(hits, now, half_life_days=90, weight=0.05):2 for h in hits:3 age = (now - h.payload["updated_at"]).days4 h.score += weight * 0.5 ** (age / half_life_days)5 return sorted(hits, key=lambda h: h.score, reverse=True)The boost is small on purpose: it breaks ties between near-duplicates without letting a fresh but irrelevant chunk beat the right answer.
is_currentflag instead of instant delete. Mark replaced chunksis_current = falseand filter on it. If an ingest goes wrong, for example a broken PDF export that empties a page, you flip the flag back instead of re-running the whole pipeline. Purge old versions after a few days.
The pipeline and its safety net
- Capture changes — CMS webhook into a queue (SQS or Kafka); bulk edits of 40 pages queue up instead of hammering the embedding API.
- Process — workers re-chunk, diff, embed in batches, upsert.
- Reconcile — a nightly job re-hashes the corpus and repairs drift from lost events.
- Measure — p95 edit-to-searchable latency, plus a daily count of chunks whose source no longer exists.
A real-life example
Scenario (illustrative numbers). A cloud-hosting provider's docs team publishes about 30 edits a day. The support bot is rebuilt every 6 hours, so after a price change at 10 AM it quotes the old price until 4 PM, and support tickets complain.
The team agrees a 5-minute freshness SLO. A queue-based worker with hash diffing achieves a p95 of 70 seconds. In the first month, a bad export publishes 12 blank pages; the is_current flag lets them restore the previous versions in two minutes. The recency boost fixes a separate issue: an old and a new pricing page were both retrieved, and the bot sometimes quoted the old one.
Follow-up questions to expect
- "Why not just delete old chunks immediately?" — You can, but a soft delete gives you an instant rollback when an ingest goes wrong. Purge on a schedule.
- "Can the recency boost cause harm?" — Yes, if it is too strong, fresh but irrelevant content wins. Keep it small and check it on the eval set.
- "What if the same fact lives in two documents?" — Freshness alone won't fix contradictions; mark one document as the source of truth or deduplicate at index time.