RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

How do you handle large-scale document updates?


One product change through the sync workerCDC event:SKU-1001 changedHash the newtext, comparestored hashSame hash:skip, zeroembedding costNew hash:delete allSKU-1001 chunksInsert newchunks withnew hashA nightly reconciliation still found 12,000 chunks from delisted products.
Replacing every chunk of a changed document is what stops an old price or policy from surviving beside the new one.

What you need to know

The building blocks

  • Stable document ids from the source system (SKU-1001, policy-leave), stored on every chunk as source_id.
  • Content hash of each document, stored on its chunks, so you can tell "changed" from "unchanged" without embedding anything.
  • Replace, not append. Delete every chunk with this source_id, then insert the new chunks. If you only upsert, and the new version has fewer chunks, the old extra chunks stay forever.
  • Change detection. Webhooks, an "updated since" query, or change data capture (CDC: a stream of row changes from the source database). Full crawls only as a periodic safety check.
  • Deletions. When a source document is deleted, its chunks must go too. This is the step teams most often forget.

The sync function

Python
import hashlibdef sync_document(col, source_id: str, text: str) -> str:    doc_hash = hashlib.sha256(text.encode()).hexdigest()    old = col.get(where={"source_id": source_id}, limit=1, include=["metadatas"])    if old["ids"] and old["metadatas"][0]["doc_hash"] == doc_hash:        return "unchanged"                                # no embedding cost    col.delete(where={"source_id": source_id})            # remove every old chunk    chunks = split(text)    col.add(ids=[f"{source_id}:{doc_hash[:8]}:{i}" for i in range(len(chunks))],            documents=chunks, embeddings=embed(chunks),            metadatas=[{"source_id": source_id, "doc_hash": doc_hash}] * len(chunks))    return f"written, {len(chunks)} chunks"

With Chroma, running this three times on one product (new, same text, edited text) returned "written", "unchanged", "written", and the collection held only the current chunks each time. Between the delete and the add, the document briefly has no chunks; if that matters, write the new chunks first under new ids, then delete the old ones.

Big changes: blue/green indexes

A new embedding model or chunking strategy changes every vector. Do not edit the live index.

  1. Build green — a new collection with the new model, filled from the source.
  2. Catch up — replay changes that happened during the build.
  3. Evaluate — run the golden set on green and compare with blue (the live one).
  4. Switch — point the app's alias or config to green.
  5. Keep blue — for a few days, so you can switch back.

Reconciliation

Every night, compare the source list of documents and hashes with what the index holds. Alert on documents missing from the index, chunks whose source no longer exists, and hash mismatches.

A real-life example

An e-commerce marketplace has 2 million products. About 30,000 change on a normal day, and 400,000 during a sale when prices and descriptions are updated. A weekly full re-crawl meant answers used week-old descriptions.

They switch to CDC from the product database into a queue. A worker calls sync_document per changed product, skipping unchanged content (most price-only changes do not touch the indexed description at all, and price is fetched live instead). Freshness lag at p95 drops from days to under 5 minutes. The nightly reconciliation found 12,000 chunks from products delisted months ago, which had been showing up in answers.

Follow-up questions to expect

  • "Why not update chunks in place?" — A changed document can split into a different number of chunks at different positions. Replacing all of them is simpler and never leaves orphans.
  • "How do you handle deletes when the vector store is slow at deleting?" — Mark chunks is_active: false, filter on it at query time, and compact later.
  • "How do you update without downtime?" — Small changes: write new then delete old. Big changes: blue/green with an alias switch.