Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to update a LangChain vector store with new documents.
What you need to know
The naive version
def add_new(vector_store, docs, splitter): chunks = splitter.split_documents(docs) return vector_store.add_documents(chunks) # returns the new idsThis is fine for a one-time load. It breaks on the second run: if you re-run it over the same folder, every chunk is added again. If a document was edited, the old chunks stay next to the new ones, and the model sees two versions of the policy.
Stable ids: the half-way fix
Most stores accept ids= in add_documents. If you build ids from the source and chunk number ("leave-policy.pdf#12"), a re-run overwrites instead of duplicating. But if an edit makes a document shorter, chunks 13–20 remain as stale leftovers.
The indexing API
1from langchain_core.indexing import index2from langchain_classic.indexes import SQLRecordManager34record_manager = SQLRecordManager(5 "pgvector/hr_policies", db_url="postgresql+psycopg://.../indexing")6record_manager.create_schema() # once78def sync(vector_store, docs, splitter):9 chunks = splitter.split_documents(docs) # each chunk keeps metadata["source"]10 return index(chunks, record_manager, vector_store,11 cleanup="incremental", source_id_key="source")1213sync(vector_store, docs, splitter)14# {'num_added': 40, 'num_updated': 0, 'num_skipped': 860, 'num_deleted': 12}The cleanup modes:
cleanup | Deletes | Use when |
|---|---|---|
None | Nothing | You only ever add |
"incremental" | Old chunks of any source that appears in this batch | Syncing only the files that changed |
"full" | Every chunk not in this batch | The batch is always the complete corpus |
"scoped_full" | Like full, but only for sources seen in this batch | Large corpora synced in parts |
index itself lives in langchain_core.indexing. SQLRecordManager is in langchain-classic in LangChain 1.x; InMemoryRecordManager in core is enough for tests. The vector store must support deleting by id.
A real-life example
An HR bot indexes 350 policy PDFs from a shared drive. A cron job ran add_documents over the whole folder every night. After three weeks the index held 21 copies of each chunk, retrieval returned the same paragraph four times, and the embedding bill was about 20 times higher than needed.
After switching to index with cleanup="incremental" and source_id_key="source", a typical night reports num_skipped: 12,400, num_added: 35, num_deleted: 18 — only the two edited PDFs were re-embedded. When HR changed the work-from-home allowance from 2 to 3 days a week, the old "2 days" chunk was deleted the same night, so the bot stopped quoting it.
Follow-up questions to expect
- "What does
source_id_keydo?" — It tellsindexwhich metadata field identifies the original document, so it knows which old chunks belong to an updated source. - "Why not
cleanup='full'every time?" — Full deletes everything not in the batch. If one night's loader fails halfway, you wipe half the index. - "How do you handle a deleted source file?" — Incremental mode never sees it, so run a periodic full or scoped-full sync, or delete by source explicitly.