RAG Systems

Course Content

RAG Systems

12 sections · 66 lessons

What is the difference between similarity search and MMR (Max Marginal Relevance)?


What you need to know

The MMR score

Text
MMR(d) = lambda_mult × sim(query, d) − (1 − lambda_mult) × max sim(d, already selected)

lambda_mult = 1 is pure similarity; 0 is pure diversity; LangChain's default is 0.5. LangChain first fetches fetch_k candidates by similarity, always takes the single most similar one first, and then applies the formula for each remaining pick.

A worked example by hand

Three candidates for "What is the refund timeline?", with lambda_mult = 0.5:

Candidatesim to querysim to d1
d1: Refund policy (current page)0.90—
d2: Refund policy (same text, older URL)0.880.97
d3: Refunds for cancelled orders0.800.40
  • Pick 1: d1, the most similar.
  • Pick 2: compute MMR for the rest. - d2: 0.5 × 0.88 − 0.5 × 0.97 = 0.44 − 0.485 = −0.045 - d3: 0.5 × 0.80 − 0.5 × 0.40 = 0.40 − 0.20 = 0.20
  • d3 wins. Plain similarity would have picked d2, a copy of what the model already has.

In code

Python
retriever = store.as_retriever(    search_type="mmr",    search_kwargs={"k": 5, "fetch_k": 30, "lambda_mult": 0.5})

When each fits

Plain similarity

  • Narrow factual question with one answer
  • Clean corpus without duplicates
  • You will rerank afterwards anyway
  • Cheapest option

MMR

  • Overlapping chunks or copied boilerplate
  • Several versions or mirrors of one page
  • Broad or comparative questions
  • You want coverage, not five phrasings

Its costs

MMR needs the candidate vectors and about fetch_k × k extra similarity calculations, which is small. The real cost is quality: it can demote a genuinely useful second passage just because it resembles the first. A reranker is often the better tool when you want the best passages; MMR is for when you want different passages.

A real-life example

An HR policy assistant indexes the same leave policy from three places: the HR portal, a PDF attached to a circular, and a copy on the intranet wiki. For "What types of leave can I take?", plain similarity returns the same "Earned leave" paragraph three times, plus two other chunks. The answer lists only earned leave and sick leave.

The long-term fix is to deduplicate at ingest (hash normalised text and keep one copy). As an immediate fix, the team switches that assistant to MMR with fetch_k = 30 and lambda_mult = 0.6. The top 5 now cover earned, sick, parental, bereavement and unpaid leave, and the answer lists all five.

Follow-up questions to expect

  • "How do you choose lambda_mult?" — Start at 0.5 to 0.7 and test on broad questions; lower values give more diversity but risk off-topic chunks.
  • "Is MMR the same as reranking?" — No. Reranking rescores each passage's relevance more accurately; MMR trades some relevance for variety using the same similarity scores.
  • "Should you fix duplicates with MMR?" — Only as a patch. Deduplicate at ingest so duplicates never reach the index.