LangChain Mastery

Course Content

LangChain Mastery

7 sections · 109 lessons

Write a function to implement a contextual compression retriever.


From 15 chunks to 900 tokensRetrieve k=15fare-rule chunksEmbeddingsFilterdrops weakones: 4 leftSmall LLMextracts therelevant linesPrompt ofabout 900tokens, not 6,000Cheap filter first, expensive extractor second.
Ordering the compressors is the whole cost story: the LLM extractor should only ever see the survivors.

What you need to know

Chunks are cut before anyone knows the question. So a retrieved 1,000-character chunk may contain one useful sentence and nine unrelated ones. Sending all of it wastes tokens and can distract the model. Compression fixes this after retrieval, using the actual question.

The compressors

CompressorWhat it doesCost
EmbeddingsFilterDrops docs whose embedding similarity to the query is below a thresholdCheap: embeddings only
EmbeddingsRedundantFilterDrops near-duplicate docsCheap
CrossEncoderRerankerScores query and doc together, keeps top nMedium: a small model
LLMChainFilterLLM says yes/no per docOne LLM call per doc
LLMChainExtractorLLM copies out only the relevant sentencesOne LLM call per doc, most expensive

The function

Python
from langchain_classic.retrievers import ContextualCompressionRetrieverfrom langchain_classic.retrievers.document_compressors import (    DocumentCompressorPipeline, EmbeddingsFilter, LLMChainExtractor)def compressed_retriever(vector_store, llm, embeddings, fetch_k: int = 15):    base = vector_store.as_retriever(search_kwargs={"k": fetch_k})    pipeline = DocumentCompressorPipeline(transformers=[        EmbeddingsFilter(embeddings=embeddings, similarity_threshold=0.75),        LLMChainExtractor.from_llm(llm),    ])    return ContextualCompressionRetriever(base_compressor=pipeline, base_retriever=base)retriever = compressed_retriever(vector_store, small_llm, embeddings)docs = retriever.invoke("What is the baggage allowance on domestic flights?")
  • Order matters — the cheap filter runs first, so the expensive LLM extractor sees maybe 4 documents instead of 15.
  • fetch_k is higher than you would normally use, because most results will be dropped.
  • Use a small, cheap model for extraction; it is a copy-out task, not reasoning.
  • The retriever is still a normal retriever, so it drops into any chain. In LangChain 1.x these classes are imported from langchain_classic.

When not to use LLM extraction

If your goal is only better ranking, a CrossEncoderReranker with top_n=5 is faster and usually enough. LLM extraction also has a risk: it can paraphrase or drop a qualifier ("except on international routes"), and the main model then answers from a weakened text.

A real-life example

A travel-booking agent answers fare-rule questions. Airline fare rules are long; a single chunk mixes baggage, cancellation, date-change and meal rules. With k=6 and no compression, the prompt was about 6,000 tokens and the model sometimes quoted the cancellation fee when asked about baggage.

With EmbeddingsFilter (threshold 0.75) and a small extractor model, the prompt dropped to about 900 tokens. Answer accuracy on their 120-question fare test set rose from 81% to 90%, but p95 latency rose by 1.1 seconds because of four extra extractor calls. They kept extraction for the fare-rules tool (accuracy matters, users wait anyway) and used a reranker only for the general FAQ path.

Follow-up questions to expect

  • "How do you pick the similarity threshold?" — Look at score distributions for known-relevant and known-irrelevant pairs on your data; there is no universal number.
  • "Does compression reduce cost?" — It reduces the main model's input tokens, but adds extractor calls; measure total cost per answer.
  • "How is this different from reranking?" — Reranking reorders and cuts whole documents; compression can also cut text inside documents.