Course Content
LangChain Mastery
7 sections · 109 lessons
Write a function to implement a contextual compression retriever.
What you need to know
Chunks are cut before anyone knows the question. So a retrieved 1,000-character chunk may contain one useful sentence and nine unrelated ones. Sending all of it wastes tokens and can distract the model. Compression fixes this after retrieval, using the actual question.
The compressors
| Compressor | What it does | Cost |
|---|---|---|
EmbeddingsFilter | Drops docs whose embedding similarity to the query is below a threshold | Cheap: embeddings only |
EmbeddingsRedundantFilter | Drops near-duplicate docs | Cheap |
CrossEncoderReranker | Scores query and doc together, keeps top n | Medium: a small model |
LLMChainFilter | LLM says yes/no per doc | One LLM call per doc |
LLMChainExtractor | LLM copies out only the relevant sentences | One LLM call per doc, most expensive |
The function
1from langchain_classic.retrievers import ContextualCompressionRetriever2from langchain_classic.retrievers.document_compressors import (3 DocumentCompressorPipeline, EmbeddingsFilter, LLMChainExtractor)45def compressed_retriever(vector_store, llm, embeddings, fetch_k: int = 15):6 base = vector_store.as_retriever(search_kwargs={"k": fetch_k})7 pipeline = DocumentCompressorPipeline(transformers=[8 EmbeddingsFilter(embeddings=embeddings, similarity_threshold=0.75),9 LLMChainExtractor.from_llm(llm),10 ])11 return ContextualCompressionRetriever(base_compressor=pipeline, base_retriever=base)1213retriever = compressed_retriever(vector_store, small_llm, embeddings)14docs = retriever.invoke("What is the baggage allowance on domestic flights?")- Order matters — the cheap filter runs first, so the expensive LLM extractor sees maybe 4 documents instead of 15.
fetch_kis higher than you would normally use, because most results will be dropped.- Use a small, cheap model for extraction; it is a copy-out task, not reasoning.
- The retriever is still a normal retriever, so it drops into any chain. In LangChain 1.x these classes are imported from
langchain_classic.
When not to use LLM extraction
If your goal is only better ranking, a CrossEncoderReranker with top_n=5 is faster and usually enough. LLM extraction also has a risk: it can paraphrase or drop a qualifier ("except on international routes"), and the main model then answers from a weakened text.
A real-life example
A travel-booking agent answers fare-rule questions. Airline fare rules are long; a single chunk mixes baggage, cancellation, date-change and meal rules. With k=6 and no compression, the prompt was about 6,000 tokens and the model sometimes quoted the cancellation fee when asked about baggage.
With EmbeddingsFilter (threshold 0.75) and a small extractor model, the prompt dropped to about 900 tokens. Answer accuracy on their 120-question fare test set rose from 81% to 90%, but p95 latency rose by 1.1 seconds because of four extra extractor calls. They kept extraction for the fare-rules tool (accuracy matters, users wait anyway) and used a reranker only for the general FAQ path.
Follow-up questions to expect
- "How do you pick the similarity threshold?" — Look at score distributions for known-relevant and known-irrelevant pairs on your data; there is no universal number.
- "Does compression reduce cost?" — It reduces the main model's input tokens, but adds extractor calls; measure total cost per answer.
- "How is this different from reranking?" — Reranking reorders and cuts whole documents; compression can also cut text inside documents.