LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to update a LangChain vector store with new documents.


What the record manager decides on a nightly syncleave.pdf#1hash unchangedskipleave.pdf#2hash changedre-embedwfh.pdf#4new chunkaddwfh.pdf#5gonefrom sourcedeleteChunkRecord managerActioncleanup=incremental removes old chunks only for sources in this batch.
Hashing each chunk turns a nightly re-index from thousands of embedding calls into a handful.

What you need to know

The naive version

Python
def add_new(vector_store, docs, splitter):    chunks = splitter.split_documents(docs)    return vector_store.add_documents(chunks)     # returns the new ids

This is fine for a one-time load. It breaks on the second run: if you re-run it over the same folder, every chunk is added again. If a document was edited, the old chunks stay next to the new ones, and the model sees two versions of the policy.

Stable ids: the half-way fix

Most stores accept ids= in add_documents. If you build ids from the source and chunk number ("leave-policy.pdf#12"), a re-run overwrites instead of duplicating. But if an edit makes a document shorter, chunks 13–20 remain as stale leftovers.

The indexing API

Python
from langchain_core.indexing import indexfrom langchain_classic.indexes import SQLRecordManagerrecord_manager = SQLRecordManager(    "pgvector/hr_policies", db_url="postgresql+psycopg://.../indexing")record_manager.create_schema()          # oncedef sync(vector_store, docs, splitter):    chunks = splitter.split_documents(docs)   # each chunk keeps metadata["source"]    return index(chunks, record_manager, vector_store,                 cleanup="incremental", source_id_key="source")sync(vector_store, docs, splitter)# {'num_added': 40, 'num_updated': 0, 'num_skipped': 860, 'num_deleted': 12}

The cleanup modes:

cleanupDeletesUse when
NoneNothingYou only ever add
"incremental"Old chunks of any source that appears in this batchSyncing only the files that changed
"full"Every chunk not in this batchThe batch is always the complete corpus
"scoped_full"Like full, but only for sources seen in this batchLarge corpora synced in parts

index itself lives in langchain_core.indexing. SQLRecordManager is in langchain-classic in LangChain 1.x; InMemoryRecordManager in core is enough for tests. The vector store must support deleting by id.

A real-life example

An HR bot indexes 350 policy PDFs from a shared drive. A cron job ran add_documents over the whole folder every night. After three weeks the index held 21 copies of each chunk, retrieval returned the same paragraph four times, and the embedding bill was about 20 times higher than needed.

After switching to index with cleanup="incremental" and source_id_key="source", a typical night reports num_skipped: 12,400, num_added: 35, num_deleted: 18 — only the two edited PDFs were re-embedded. When HR changed the work-from-home allowance from 2 to 3 days a week, the old "2 days" chunk was deleted the same night, so the bot stopped quoting it.

Follow-up questions to expect

  • "What does source_id_key do?" — It tells index which metadata field identifies the original document, so it knows which old chunks belong to an updated source.
  • "Why not cleanup='full' every time?" — Full deletes everything not in the batch. If one night's loader fails halfway, you wipe half the index.
  • "How do you handle a deleted source file?" — Incremental mode never sees it, so run a periodic full or scoped-full sync, or delete by source explicitly.