LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

How do you optimize vector store retrieval in LangChain?


Recall at 5 on 150 claims questions, one change at a timeBaseline: fixed 1,000-char chunks — 0.62Split on policy headings — 0.74Fetch 25, rerank to 5 — 0.86Filter by policy type first — 0.90Tune HNSW ef_search: 45 to 18 ms
Quality levers come first and are measured on one fixed set; index tuning buys speed only after recall is right.

What you need to know

"Optimise" can mean better quality (the right chunk is found) or better speed and cost. Quality usually matters more, and speed tuning is only worth doing once quality is good.

Quality levers, in order of impact

  1. Chunking — 500–1,000 characters with 10–15% overlap, split on headings or paragraphs (MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter). Keep the heading in the chunk text or metadata.
  2. Retrieve wide, then rerank — fetch 20–30 candidates with the vector store, then use a cross-encoder or a hosted rerank model to pick the best 4–5.
  3. Metadata filters — search only the product, language or year that applies.
  4. Hybrid search — combine BM25 keyword search with vectors so ids, codes and names match.
  5. Query rewriting — rewrite vague or follow-up questions into a clear standalone query (or use MultiQueryRetriever) before searching.
Python
from langchain_classic.retrievers import ContextualCompressionRetrieverfrom langchain_classic.retrievers.document_compressors import CrossEncoderRerankerfrom langchain_community.cross_encoders import HuggingFaceCrossEncoderbase = vector_store.as_retriever(search_kwargs={"k": 25})reranker = CrossEncoderReranker(    model=HuggingFaceCrossEncoder(model_name="BAAI/bge-reranker-base"), top_n=5)retriever = ContextualCompressionRetriever(base_compressor=reranker, base_retriever=base)

Speed and cost levers

  • ANN index settings — for HNSW, ef_search and m trade recall for latency; for IVF, nprobe. This is where milliseconds come from.
  • Smaller vectors — fewer dimensions or int8 quantisation cut memory and speed up search.
  • Cache query embeddings for repeated questions, and use CacheBackedEmbeddings so re-indexing never pays twice for the same chunk.
  • Keep k small — every extra chunk adds prompt tokens and model latency.

Measure

Use recall@k (was the right chunk in the top k?) and MRR (how high was it?) on 50–200 real questions. Put the set in LangSmith so every change is a comparable experiment.

A real-life example

An insurance company's claims assistant had recall@5 of 0.62 on 150 labelled questions: the right clause was missing from the top 5 for 38% of questions.

  • Header-aware chunking (one policy clause per chunk): recall@5 rose to 0.74.
  • Retrieve 25, rerank to 5 with a cross-encoder: 0.86, at a cost of about 120 ms extra latency.
  • Metadata filter on policy type (health, motor, travel): 0.90, and search got faster because the filtered set was smaller.

They then lowered HNSW ef_search until recall started to drop, which cut vector-search time from 45 ms to 18 ms at the same recall. All numbers came from the same test set, so each step was a fair comparison.

Follow-up questions to expect

  • "Is a bigger embedding model the first fix?" — Rarely. Chunking and reranking usually give more for less cost; test a new model only after those.
  • "What does reranking cost?" — One cross-encoder pass per candidate, typically tens to low hundreds of milliseconds for 25 candidates; hosted rerank APIs charge per call.
  • "How do you know the optimisation helped end to end?" — Also track answer-level metrics such as faithfulness and correctness on the same set, not only retrieval metrics.