Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
Scenario – 3: Redundant Retrieval Results
What you need to know
The scenario: the top 5 chunks are near-copies of each other — often from the same long document or from overlapping chunks — and the answer misses important points.
Why similarity search repeats itself
Each chunk is scored against the query independently. Nothing in plain top-k asks, "have I already got this information?" If a document is split with 50% overlap, neighbouring chunks share half their text and score almost the same. The context fills with one fact repeated.
1import numpy as np23def mmr(query_vec, cand_vecs, k=5, lam=0.6):4 rel = cand_vecs @ query_vec # vectors are normalised5 chosen = [int(rel.argmax())]6 while len(chosen) < min(k, len(cand_vecs)):7 sim_to_chosen = (cand_vecs @ cand_vecs[chosen].T).max(axis=1)8 score = lam * rel - (1 - lam) * sim_to_chosen9 score[chosen] = -np.inf10 chosen.append(int(score.argmax()))11 return chosenFixes at three stages
| Stage | Fix | Note |
|---|---|---|
| Index time | Exact hash and MinHash deduplication, strip boilerplate | Removes the cause |
| Index time | Overlap of about 10–15% instead of 50% | High overlap creates near-duplicates by design |
| Retrieval | MMR with lambda around 0.5–0.7 | Tune on the golden set |
| Retrieval | At most 2 chunks per source document | Stops one long document owning the top 5 |
| After retrieval | Retrieve 50, rerank, compress to relevant sentences | Frees room for distinct sources |
The trade-off
Push diversity too hard and you pull in chunks that are different but barely relevant, which dilutes the answer. Measure both distinct source documents in the final context and coverage of the gold answer's supporting facts, and pick lambda where coverage peaks.
A real-life example
Scenario, numbers made up. A pharma company's medical-information assistant indexes product monographs with 50% chunk overlap. Asked about side effects and drug interactions for one product, the top 5 are five overlapping chunks from the side-effects section; the interactions table never appears.
The team re-chunks with 12% overlap, adds a two-chunks-per-section cap and MMR with lambda 0.6. Distinct sections in the final context rise from 1.4 to 3.6 on average, and coverage of supporting facts on a 120-question golden set rises from 58% to 84%. At lambda 0.3, coverage fell again because loosely related sections crowded in — which is why they tuned it rather than guessing.
Follow-up questions to expect
- "Does a reranker fix redundancy?" — No. A reranker scores each chunk independently too, so duplicates still rank together; apply MMR or a per-source cap after reranking.
- "Why keep any overlap at all?" — A little overlap protects sentences cut at chunk edges; the problem is large overlap.
- "How do you choose lambda?" — Sweep it on the golden set and pick the value where supporting-fact coverage is highest.