Course Content
RAG Systems
12 sections · 66 lessons
What is the difference between similarity search and MMR (Max Marginal Relevance)?
What you need to know
The MMR score
MMR(d) = lambda_mult × sim(query, d) − (1 − lambda_mult) × max sim(d, already selected)lambda_mult = 1 is pure similarity; 0 is pure diversity; LangChain's default is 0.5. LangChain first fetches fetch_k candidates by similarity, always takes the single most similar one first, and then applies the formula for each remaining pick.
A worked example by hand
Three candidates for "What is the refund timeline?", with lambda_mult = 0.5:
| Candidate | sim to query | sim to d1 |
|---|---|---|
| d1: Refund policy (current page) | 0.90 | — |
| d2: Refund policy (same text, older URL) | 0.88 | 0.97 |
| d3: Refunds for cancelled orders | 0.80 | 0.40 |
- Pick 1: d1, the most similar.
- Pick 2: compute MMR for the rest. - d2: 0.5 × 0.88 − 0.5 × 0.97 = 0.44 − 0.485 = −0.045 - d3: 0.5 × 0.80 − 0.5 × 0.40 = 0.40 − 0.20 = 0.20
- d3 wins. Plain similarity would have picked d2, a copy of what the model already has.
In code
retriever = store.as_retriever( search_type="mmr", search_kwargs={"k": 5, "fetch_k": 30, "lambda_mult": 0.5})When each fits
Plain similarity
- Narrow factual question with one answer
- Clean corpus without duplicates
- You will rerank afterwards anyway
- Cheapest option
MMR
- Overlapping chunks or copied boilerplate
- Several versions or mirrors of one page
- Broad or comparative questions
- You want coverage, not five phrasings
Its costs
MMR needs the candidate vectors and about fetch_k × k extra similarity calculations, which is small. The real cost is quality: it can demote a genuinely useful second passage just because it resembles the first. A reranker is often the better tool when you want the best passages; MMR is for when you want different passages.
A real-life example
An HR policy assistant indexes the same leave policy from three places: the HR portal, a PDF attached to a circular, and a copy on the intranet wiki. For "What types of leave can I take?", plain similarity returns the same "Earned leave" paragraph three times, plus two other chunks. The answer lists only earned leave and sick leave.
The long-term fix is to deduplicate at ingest (hash normalised text and keep one copy). As an immediate fix, the team switches that assistant to MMR with fetch_k = 30 and lambda_mult = 0.6. The top 5 now cover earned, sick, parental, bereavement and unpaid leave, and the answer lists all five.
Follow-up questions to expect
- "How do you choose
lambda_mult?" — Start at 0.5 to 0.7 and test on broad questions; lower values give more diversity but risk off-topic chunks. - "Is MMR the same as reranking?" — No. Reranking rescores each passage's relevance more accurately; MMR trades some relevance for variety using the same similarity scores.
- "Should you fix duplicates with MMR?" — Only as a patch. Deduplicate at ingest so duplicates never reach the index.