Course Content
Advanced RAG
3 sections · 38 lessons
How can you update a RAG index incrementally without rebuilding it?
What you need to know
Why not rebuild nightly
A full rebuild re-embeds everything, which costs money and time, and the index is stale for up to a day. For a policy that changed at 10 a.m., "the assistant will know tomorrow" is often unacceptable. Most days, only a tiny fraction of documents change.
The building blocks
- Stable IDs —
{doc_id}:{chunk_no}or an ID derived from the section anchor. The same chunk must get the same ID on every run. - Content hash — a hash of the chunk text. Same hash means nothing to do.
- Metadata —
doc_version,source_updated_at,effective_from,active, access labels. - Change detection — change-data-capture (CDC) from a database, webhooks from a document system, or a query for rows with
updated_atafter the last run.
The update logic
1import hashlib23def h(text):4 return hashlib.sha256(text.encode()).hexdigest()[:12]56def plan_update(doc_id, new_chunks, stored):7 """stored: {chunk_key: content_hash} currently in the index for this doc."""8 upserts, keep = [], set()9 for i, text in enumerate(new_chunks):10 key = f"{doc_id}:{i}"11 keep.add(key)12 if stored.get(key) != h(text):13 upserts.append(key) # new or changed: re-embed14 deletes = [k for k in stored if k not in keep] # chunk no longer exists15 return upserts, deletes1617old = ["Scope: tablet line 2.", "Clean with 70% IPA.", "Record in log A."]18stored = {f"SOP-114:{i}": h(t) for i, t in enumerate(old)}19new = ["Scope: tablet line 2.", "Clean with 70% IPA, then dry 10 min."]20print(plan_update("SOP-114", new, stored)) # (['SOP-114:1'], ['SOP-114:2'])Chunk 0 is unchanged and skipped. Chunk 1 changed and is re-embedded. Chunk 2 disappeared and is deleted. Only one embedding call is needed instead of three.
One weakness to mention: position-based keys shift when a paragraph is inserted near the top, so every later chunk looks "changed". Keying chunks by section heading or by content hash avoids most of that churn.
Deletes and consistency
- Soft delete first. Set
active = falseand filter on it in every query; hard-delete later in a compaction job. This avoids a moment where half of a document's chunks are new and half are old. - Write the new version, then switch. Insert the new chunks with
doc_version = 5, then mark version 4 inactive in one step. - Clean up caches. A semantic cache holding an answer built from the old version must be invalidated for that document.
Operational notes
- HNSW indexes handle inserts well but degrade under heavy delete churn: deleted points linger as tombstones and recall slowly drifts. Schedule compaction or periodic re-indexing.
- Derived artefacts — hypothetical questions, summaries, contextual headers, knowledge-graph edges — must be regenerated for changed chunks too.
- Switching the embedding model is not an update. The vector space changes, so it is a full re-embedding and migration.
A real-life example
A pharma company's regulatory-document search covers 60,000 SOPs, guidelines and templates. Around 150 documents change each day as SOPs are revised. A superseded SOP must never be served: an operator following the old cleaning procedure is a compliance deviation.
The pipeline:
- The document management system sends a webhook when an SOP reaches "Effective" status.
- A worker fetches it, re-chunks by section, and hashes each chunk. On a typical revision, 2–3 sections of 20 change.
- Changed chunks are embedded and written with the new
doc_versionandeffective_fromdate. - In the same transaction, all chunks of the previous version get
active = false. - The semantic cache entries that cited that SOP are deleted.
Every query filters on active = true. A weekly job hard-deletes inactive chunks older than 90 days (kept that long for audit questions like "what did version 4 say?", served from a separate archive index). Daily embedding cost falls to a small fraction of a full rebuild, and a revised SOP is searchable within minutes of approval.
Follow-up questions to expect
- "How do you handle a document that's deleted at the source?" — The change feed must emit deletes too; soft-delete all its chunks immediately. Also run a periodic reconciliation that compares source IDs with index IDs to catch missed events.
- "What about access-control changes?" — Treat them as metadata updates: update the ACL field on the chunks without re-embedding. Permission changes should apply within minutes.
- "How do you know the index is in sync?" — Track lag between source update and index update, and count documents in the source versus active documents in the index.