Course Content
LangChain Mastery
7 sections · 109 lessons
How do you handle large-scale vector stores in LangChain?
What you need to know
What breaks at scale
- Memory — 10 million vectors × 1,536 dimensions × 4 bytes is about 61 GB of raw floats, before index overhead. An in-process store in each web worker cannot hold that.
- Search time — exact search compares the query with every vector. You need an approximate index.
- Indexing cost — re-embedding everything after a change takes hours and real money.
- Freshness — updates must flow in without rebuilding the index.
Index choices
| Index | Idea | Trade-off |
|---|---|---|
| HNSW | A layered graph you walk towards the query | Fast and accurate; uses a lot of memory |
| IVF | Cluster vectors, search only the nearest clusters (nprobe) | Less memory; recall depends on nprobe |
| PQ / scalar quantisation | Store compressed vectors (e.g. int8) | 4× or more memory saving; a few points of recall lost |
The practical checklist
- Use a server store — pgvector, Qdrant, Milvus, Weaviate, Pinecone, Elasticsearch/OpenSearch. LangChain's integration just connects to it.
- Shrink vectors — fewer dimensions (if the model supports it) or quantisation.
- Partition — by tenant, region or time, so most queries hit one partition.
- Filter early — metadata filters that the database applies inside the index.
- Index offline and incrementally — batch
embed_documentscalls (hundreds of texts per call), run as a job, and use the indexing API with a record manager so only changed chunks are re-embedded. - Monitor — p95 search latency, recall on a fixed test set, index size and embedding spend.
LangChain's role here is small on purpose: as_retriever() over the database client. The scaling work is database work.
A real-life example
A legal-research startup indexes 40 million judgment paragraphs from Indian courts. The prototype used FAISS in each API pod; memory reached 70 GB per pod and startup took 20 minutes.
The production design: Qdrant with three shards and one replica each, int8 scalar quantisation (memory down about 4×, recall@10 down from 0.93 to 0.91 on their test set), and partitioning by court. Queries filtered to one High Court search about 5% of the data. New judgments arrive daily and are indexed by a nightly job using a record manager — about 30,000 new paragraphs a night instead of 40 million. p95 search latency is 35 ms.
Follow-up questions to expect
- "When is pgvector enough?" — Up to several million vectors with HNSW when you already run Postgres and want SQL joins; beyond that, or for heavy write loads, a dedicated vector database is easier to scale.
- "How do you change the embedding model at this scale?" — Build a new index in parallel (blue-green), backfill in batches, compare recall, then switch reads.
- "Where does latency go?" — Often the query embedding API call, not the vector search; cache query embeddings and keep the embedding service close.