Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
You indexed 10,000 internal policy PDFs but many have repeated boilerplate sections. Retrieval keeps returning these duplicates instead of actual answers. How do you detect and de-duplicate chunks at index time?
What you need to know
Why boilerplate wins retrieval
A confidentiality notice appears in all 10,000 PDFs, so the index holds 10,000 near-identical chunks. Many questions share words with it ("policy", "employee", "applicable"), and each copy scores about the same. The top 5 fills with copies of text that answers nothing.
Step 1: strip boilerplate at parse time
Count how many documents each normalised text block (a paragraph or line) appears in. Blocks in more than about 20% of documents are almost always headers, footers or template text. Remove them before chunking. This removes the cause and shrinks the index.
Step 2: deduplicate in tiers
| Tier | Method | Catches | Cost |
|---|---|---|---|
| Exact | Normalise whitespace and case, SHA-256 hash | Identical copies | Almost free |
| Near-duplicate | MinHash plus LSH on word shingles, Jaccard about 0.85 | Same clause with a changed date or name | Cheap, scales to millions |
| Semantic | Embedding cosine above about 0.95 | Reworded copies | Expensive; run on what remains |
1from datasketch import MinHash, MinHashLSH23def shingles(text, n=5):4 words = text.lower().split()5 return {" ".join(words[i:i + n]) for i in range(max(1, len(words) - n + 1))}67lsh = MinHashLSH(threshold=0.85, num_perm=128)8canonical_of = {} # chunk_id -> canonical chunk_id9for chunk in chunks:10 m = MinHash(num_perm=128)11 for s in shingles(chunk.text):12 m.update(s.encode("utf8"))13 match = lsh.query(m)14 if match:15 canonical_of[chunk.id] = match[0] # keep provenance, skip indexing16 else:17 lsh.insert(chunk.id, m)18 canonical_of[chunk.id] = chunk.idKeep provenance
A clause that really appears in 40 policies is legitimately part of all 40. Index it once, but keep the canonical-to-sources map, so when a user asks about Policy B the answer can cite Policy B, not "Policy A, where we happened to keep the copy". Filters by document or department must still work, so store the list of source documents on the canonical chunk.
Step 3: a query-time safety net
Use MMR (maximal marginal relevance), which picks results that are relevant but different from those already chosen, with a balance of about 0.5, or simply cap results at two chunks per source document.
A real-life example
Scenario, numbers made up. A manufacturing company indexes 10,000 policy PDFs, about 1.8M chunks. Asked "How many days of bereavement leave do I get?", the assistant's top 5 are four copies of a legal disclaimer and one page header.
A frequency pass finds 23 blocks that appear in more than 20% of documents; stripping them removes 31% of all chunks. Exact hashing removes another 9%, and MinHash finds 6% more, mostly the same clause across regional versions with different dates. The index shrinks to about 1M chunks. The average number of distinct source documents in the final context rises from 1.6 to 4.2, and accuracy on a 200-question golden set rises from 64% to 81%.
Follow-up questions to expect
- "What if the 'duplicate' differs in one important number?" — That is why the near-duplicate threshold must be tested on real pairs; for policies, treat differing numbers or dates as different chunks, or keep both versions with effective dates as metadata.
- "Why not only use MMR at query time?" — It hides the symptom but keeps a bloated index, wastes candidate slots on copies and costs more; fix ingestion and keep MMR as a safety net.
- "How do you prove it worked?" — Index size, distinct sources in the final context, and golden-set accuracy before and after.