Scenario-Based AI Engineering Questions

Course Content

Scenario-Based AI Engineering Questions

26 sections · 146 lessons

Scenario – 3: Redundant Retrieval Results


What you need to know

The scenario: the top 5 chunks are near-copies of each other — often from the same long document or from overlapping chunks — and the answer misses important points.

Why similarity search repeats itself

Each chunk is scored against the query independently. Nothing in plain top-k asks, "have I already got this information?" If a document is split with 50% overlap, neighbouring chunks share half their text and score almost the same. The context fills with one fact repeated.

Python
import numpy as npdef mmr(query_vec, cand_vecs, k=5, lam=0.6):    rel = cand_vecs @ query_vec                            # vectors are normalised    chosen = [int(rel.argmax())]    while len(chosen) < min(k, len(cand_vecs)):        sim_to_chosen = (cand_vecs @ cand_vecs[chosen].T).max(axis=1)        score = lam * rel - (1 - lam) * sim_to_chosen        score[chosen] = -np.inf        chosen.append(int(score.argmax()))    return chosen

Fixes at three stages

StageFixNote
Index timeExact hash and MinHash deduplication, strip boilerplateRemoves the cause
Index timeOverlap of about 10–15% instead of 50%High overlap creates near-duplicates by design
RetrievalMMR with lambda around 0.5–0.7Tune on the golden set
RetrievalAt most 2 chunks per source documentStops one long document owning the top 5
After retrievalRetrieve 50, rerank, compress to relevant sentencesFrees room for distinct sources

The trade-off

Push diversity too hard and you pull in chunks that are different but barely relevant, which dilutes the answer. Measure both distinct source documents in the final context and coverage of the gold answer's supporting facts, and pick lambda where coverage peaks.

A real-life example

Scenario, numbers made up. A pharma company's medical-information assistant indexes product monographs with 50% chunk overlap. Asked about side effects and drug interactions for one product, the top 5 are five overlapping chunks from the side-effects section; the interactions table never appears.

The team re-chunks with 12% overlap, adds a two-chunks-per-section cap and MMR with lambda 0.6. Distinct sections in the final context rise from 1.4 to 3.6 on average, and coverage of supporting facts on a 120-question golden set rises from 58% to 84%. At lambda 0.3, coverage fell again because loosely related sections crowded in — which is why they tuned it rather than guessing.

Follow-up questions to expect

  • "Does a reranker fix redundancy?" — No. A reranker scores each chunk independently too, so duplicates still rank together; apply MMR or a per-source cap after reranking.
  • "Why keep any overlap at all?" — A little overlap protects sentences cut at chunk edges; the problem is large overlap.
  • "How do you choose lambda?" — Sweep it on the golden set and pick the value where supporting-fact coverage is highest.