LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you handle large-scale vector stores in LangChain?


What changes past ten million vectors10M vectorsAbout 61 GB of raw floatsExact searchtoo slow: use ANNQuantise to cut memory 4xPartition by tenant or courtIndex nightly, incrementallyReplicas and aserver database
At scale LangChain stays a thin client; memory, index type and partitioning decide cost and latency.

What you need to know

What breaks at scale

  • Memory — 10 million vectors × 1,536 dimensions × 4 bytes is about 61 GB of raw floats, before index overhead. An in-process store in each web worker cannot hold that.
  • Search time — exact search compares the query with every vector. You need an approximate index.
  • Indexing cost — re-embedding everything after a change takes hours and real money.
  • Freshness — updates must flow in without rebuilding the index.

Index choices

IndexIdeaTrade-off
HNSWA layered graph you walk towards the queryFast and accurate; uses a lot of memory
IVFCluster vectors, search only the nearest clusters (nprobe)Less memory; recall depends on nprobe
PQ / scalar quantisationStore compressed vectors (e.g. int8)4× or more memory saving; a few points of recall lost

The practical checklist

  1. Use a server store — pgvector, Qdrant, Milvus, Weaviate, Pinecone, Elasticsearch/OpenSearch. LangChain's integration just connects to it.
  2. Shrink vectors — fewer dimensions (if the model supports it) or quantisation.
  3. Partition — by tenant, region or time, so most queries hit one partition.
  4. Filter early — metadata filters that the database applies inside the index.
  5. Index offline and incrementally — batch embed_documents calls (hundreds of texts per call), run as a job, and use the indexing API with a record manager so only changed chunks are re-embedded.
  6. Monitor — p95 search latency, recall on a fixed test set, index size and embedding spend.

LangChain's role here is small on purpose: as_retriever() over the database client. The scaling work is database work.

A real-life example

A legal-research startup indexes 40 million judgment paragraphs from Indian courts. The prototype used FAISS in each API pod; memory reached 70 GB per pod and startup took 20 minutes.

The production design: Qdrant with three shards and one replica each, int8 scalar quantisation (memory down about 4×, recall@10 down from 0.93 to 0.91 on their test set), and partitioning by court. Queries filtered to one High Court search about 5% of the data. New judgments arrive daily and are indexed by a nightly job using a record manager — about 30,000 new paragraphs a night instead of 40 million. p95 search latency is 35 ms.

Follow-up questions to expect

  • "When is pgvector enough?" — Up to several million vectors with HNSW when you already run Postgres and want SQL joins; beyond that, or for heavy write loads, a dedicated vector database is easier to scale.
  • "How do you change the embedding model at this scale?" — Build a new index in parallel (blue-green), backfill in batches, compare recall, then switch reads.
  • "Where does latency go?" — Often the query embedding API call, not the vector search; cache query embeddings and keep the embedding service close.