Advanced RAG

Course Content

Advanced RAG

3 sections · 38 lessons

How can you update a RAG index incrementally without rebuilding it?


What you need to know

Why not rebuild nightly

A full rebuild re-embeds everything, which costs money and time, and the index is stale for up to a day. For a policy that changed at 10 a.m., "the assistant will know tomorrow" is often unacceptable. Most days, only a tiny fraction of documents change.

The building blocks

  • Stable IDs — {doc_id}:{chunk_no} or an ID derived from the section anchor. The same chunk must get the same ID on every run.
  • Content hash — a hash of the chunk text. Same hash means nothing to do.
  • Metadata — doc_version, source_updated_at, effective_from, active, access labels.
  • Change detection — change-data-capture (CDC) from a database, webhooks from a document system, or a query for rows with updated_at after the last run.

The update logic

Python
import hashlibdef h(text):    return hashlib.sha256(text.encode()).hexdigest()[:12]def plan_update(doc_id, new_chunks, stored):    """stored: {chunk_key: content_hash} currently in the index for this doc."""    upserts, keep = [], set()    for i, text in enumerate(new_chunks):        key = f"{doc_id}:{i}"        keep.add(key)        if stored.get(key) != h(text):            upserts.append(key)                 # new or changed: re-embed    deletes = [k for k in stored if k not in keep]   # chunk no longer exists    return upserts, deletesold = ["Scope: tablet line 2.", "Clean with 70% IPA.", "Record in log A."]stored = {f"SOP-114:{i}": h(t) for i, t in enumerate(old)}new = ["Scope: tablet line 2.", "Clean with 70% IPA, then dry 10 min."]print(plan_update("SOP-114", new, stored))   # (['SOP-114:1'], ['SOP-114:2'])

Chunk 0 is unchanged and skipped. Chunk 1 changed and is re-embedded. Chunk 2 disappeared and is deleted. Only one embedding call is needed instead of three.

One weakness to mention: position-based keys shift when a paragraph is inserted near the top, so every later chunk looks "changed". Keying chunks by section heading or by content hash avoids most of that churn.

Deletes and consistency

  • Soft delete first. Set active = false and filter on it in every query; hard-delete later in a compaction job. This avoids a moment where half of a document's chunks are new and half are old.
  • Write the new version, then switch. Insert the new chunks with doc_version = 5, then mark version 4 inactive in one step.
  • Clean up caches. A semantic cache holding an answer built from the old version must be invalidated for that document.

Operational notes

  • HNSW indexes handle inserts well but degrade under heavy delete churn: deleted points linger as tombstones and recall slowly drifts. Schedule compaction or periodic re-indexing.
  • Derived artefacts — hypothetical questions, summaries, contextual headers, knowledge-graph edges — must be regenerated for changed chunks too.
  • Switching the embedding model is not an update. The vector space changes, so it is a full re-embedding and migration.

A real-life example

A pharma company's regulatory-document search covers 60,000 SOPs, guidelines and templates. Around 150 documents change each day as SOPs are revised. A superseded SOP must never be served: an operator following the old cleaning procedure is a compliance deviation.

The pipeline:

  1. The document management system sends a webhook when an SOP reaches "Effective" status.
  2. A worker fetches it, re-chunks by section, and hashes each chunk. On a typical revision, 2–3 sections of 20 change.
  3. Changed chunks are embedded and written with the new doc_version and effective_from date.
  4. In the same transaction, all chunks of the previous version get active = false.
  5. The semantic cache entries that cited that SOP are deleted.

Every query filters on active = true. A weekly job hard-deletes inactive chunks older than 90 days (kept that long for audit questions like "what did version 4 say?", served from a separate archive index). Daily embedding cost falls to a small fraction of a full rebuild, and a revised SOP is searchable within minutes of approval.

Follow-up questions to expect

  • "How do you handle a document that's deleted at the source?" — The change feed must emit deletes too; soft-delete all its chunks immediately. Also run a periodic reconciliation that compares source IDs with index IDs to catch missed events.
  • "What about access-control changes?" — Treat them as metadata updates: update the ACL field on the chunks without re-embedding. Permission changes should apply within minutes.
  • "How do you know the index is in sync?" — Track lag between source update and index update, and count documents in the source versus active documents in the index.