Course Content
Scenario-Based AI Engineering Questions
26 sections · 146 lessons
A malicious user manipulates retrieval rankings by repeatedly uploading keyword-stuffed documents. How do you defend vector search systems against retrieval poisoning attacks?
What you need to know
What stuffing looks like
A stuffed document repeats target phrases ("refund policy refund process how to get refund...") so its embedding and its keyword score sit close to many queries. It has a signature:
1from collections import Counter23def stuffing_features(text: str) -> dict:4 words = text.lower().split()5 counts = Counter(words)6 top_share = sum(c for _, c in counts.most_common(5)) / max(len(words), 1)7 trigrams = [" ".join(words[i:i + 3]) for i in range(len(words) - 2)]8 repeat_trigrams = 1 - len(set(trigrams)) / max(len(trigrams), 1)9 return {"unique_ratio": len(counts) / max(len(words), 1),10 "top5_share": top_share,11 "repeated_trigrams": repeat_trigrams}Normal prose has a high share of unique words and few repeated three-word phrases. Compare each upload's features with your corpus and quarantine strong outliers for review before embedding. A second signal is hubness: an embedding that is unusually close to many unrelated query clusters.
Rank with trust, not similarity alone
| Defence | Why it helps |
|---|---|
| Provenance on every chunk | Uploader, source, time and trust tier make ranking and cleanup possible |
| Upload quotas per user | One account cannot add 10,000 documents to a shared corpus unreviewed |
| Source-trust prior | Official docs above wiki above user uploads, blended into the score |
| Cross-encoder reranker | Reads query and passage together, so simple repetition helps much less |
| Diversity (MMR) and per-source caps | One uploader cannot fill the top-K |
A stuffed document can win on cosine similarity; it should not win the blended score.
Better attackers write fluent text aimed at specific questions rather than repeating keywords. Research such as PoisonedRAG has shown that a few crafted passages can steer answers, which is why provenance and trust tiers matter more than any single content filter.
Detect and clean up
- Score at ingestion — stuffing features and hubness; quarantine outliers.
- Rank with trust — blended score, rerank, per-source caps.
- Watch fan-out — alert when a document suddenly appears in the top-K for a wide range of unrelated queries, and auto-quarantine above a threshold.
- Purge by provenance — remove everything a bad uploader added in one operation, and re-run affected queries to check answers.
A real-life example
Scenario, numbers made up. A marketplace's help assistant searches official help pages plus seller-uploaded guides. One seller uploads 3,000 short documents stuffed with "refund", "cancel order" and "customer care number", each containing a fake phone number. Within a day, 7 of the top 8 results for refund questions are from that seller, and the assistant starts quoting the scam number.
Fan-out monitoring flags the documents because they rank for 400 unrelated query clusters. The team purges everything from that uploader using provenance, adds a trust prior that ranks official pages first, caps results at 2 per uploader, and rejects uploads with a top-5-word share above the corpus 99th percentile. New seller uploads now go through a review queue once a seller passes 50 documents a day.
Follow-up questions to expect
- "Is a cross-encoder enough?" — No. It makes simple stuffing much less effective, but crafted, fluent passages can still rank; you need trust tiers and provenance too.
- "How would you tell the user about lower-trust sources?" — Show the source type on each citation, and prefer official sources for policy answers such as refunds or phone numbers.
- "What about poisoning through a trusted source?" — Treat any editable source as lower trust than reviewed content, and version official pages so a bad edit can be rolled back.