Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Your RAG system indexes product documentation updated 20+ times a day. Users get outdated answers within hours. How do you keep the vector index fresh without re-embedding the entire corpus every time?


One docs edit, a handful of embeddingsCMS webhookon publishQueue absorbsbursts and rate limitsRe-chunk,comparetext hashesUpsert 3 to 6changed chunksDelete orphansby doc id prefixA nightly job re-hashes everything to catch lost webhooks.
Deleting removed chunks is the step teams forget, and it is why bots keep quoting prices that no longer exist.

What you need to know

This version of the question is really about targets and trade-offs: how fresh is fresh enough, what it costs, and how to protect answers while the index catches up.

The target decides the architecture

TargetDesignCost and complexity
Under a minuteEvent-driven: webhook, stream (Kafka), always-on workersHighest; needs on-call for the pipeline
About 5 to 10 minutesMicro-batch: a job every few minutes processes changed documentsModerate; simple to operate
DailyNightly incremental jobCheapest; fine for slow-changing content

For product docs edited 20+ times a day, a 5-minute target is usually enough, and a micro-batch or a small queue-based worker meets it.

Incremental indexing, briefly

Each chunk has a stable id (doc_id plus heading or position) and a stored hash of its text. An update re-chunks the document, compares hashes, upserts only changed chunks and deletes orphans. Twenty edits a day become a few hundred embeddings, not a few million. Use content-anchored splitting, on headings, so a one-word change near the top does not shift every chunk id and re-embed the whole document.

Protecting answers during change

  • Recency boost at rank time. Store updated_at on every chunk. When an old and a new chunk both match, add a small boost to the newer one:
Python
def rerank_with_recency(hits, now, half_life_days=90, weight=0.05):    for h in hits:        age = (now - h.payload["updated_at"]).days        h.score += weight * 0.5 ** (age / half_life_days)    return sorted(hits, key=lambda h: h.score, reverse=True)

The boost is small on purpose: it breaks ties between near-duplicates without letting a fresh but irrelevant chunk beat the right answer.

  • is_current flag instead of instant delete. Mark replaced chunks is_current = false and filter on it. If an ingest goes wrong, for example a broken PDF export that empties a page, you flip the flag back instead of re-running the whole pipeline. Purge old versions after a few days.

The pipeline and its safety net

  1. Capture changes — CMS webhook into a queue (SQS or Kafka); bulk edits of 40 pages queue up instead of hammering the embedding API.
  2. Process — workers re-chunk, diff, embed in batches, upsert.
  3. Reconcile — a nightly job re-hashes the corpus and repairs drift from lost events.
  4. Measure — p95 edit-to-searchable latency, plus a daily count of chunks whose source no longer exists.

A real-life example

Scenario (illustrative numbers). A cloud-hosting provider's docs team publishes about 30 edits a day. The support bot is rebuilt every 6 hours, so after a price change at 10 AM it quotes the old price until 4 PM, and support tickets complain.

The team agrees a 5-minute freshness SLO. A queue-based worker with hash diffing achieves a p95 of 70 seconds. In the first month, a bad export publishes 12 blank pages; the is_current flag lets them restore the previous versions in two minutes. The recency boost fixes a separate issue: an old and a new pricing page were both retrieved, and the bot sometimes quoted the old one.

Follow-up questions to expect

  • "Why not just delete old chunks immediately?" — You can, but a soft delete gives you an instant rollback when an ingest goes wrong. Purge on a schedule.
  • "Can the recency boost cause harm?" — Yes, if it is too strong, fresh but irrelevant content wins. Keep it small and check it on the eval set.
  • "What if the same fact lives in two documents?" — Freshness alone won't fix contradictions; mark one document as the source of truth or deduplicate at index time.