Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is data poisoning, and how can it impact models?
What you need to know
The variants
- Label flipping: wrong labels on some examples, which lowers accuracy.
- Clean-label poisoning: correctly labelled but carefully crafted examples that shift the decision boundary for one target.
- Backdoors: a trigger (a rare phrase, a pixel pattern) mapped to an attacker's output. Everything else stays normal.
- Fine-tuning poisoning: a few hundred harmful examples in a fine-tuning set can undo safety training.
- RAG corpus poisoning: a page or document written so it is retrieved for a target question and pushes a false answer or an injection. Research such as PoisonedRAG showed a handful of crafted documents in a large corpus can control answers for chosen questions.
- Web-scale poisoning: researchers have shown that buying expired domains that appear in public training datasets lets an attacker change what is downloaded.
Why it is hard to catch
Aggregate eval metrics look fine by design. A backdoor that fires on one phrase changes almost nothing in average accuracy. You need targeted checks.
Defences, in order of payoff
- Provenance and write control: know where every training and indexed document came from; require authentication and review for user-contributed content entering a corpus or vector store.
- Curation and screening: deduplicate; flag outliers by loss or embedding distance; review data from low-trust sources.
- A trusted holdout: evaluate every candidate model on a clean, internally controlled set before promotion.
- Behavioural evals and output guardrails: assume screening missed something.
- For RAG: treat the index as a production database, with access control, an audit trail, per-document source metadata, and the ability to remove a document and everything derived from it.
A real-life example
A company's internal HR assistant answers policy questions from a wiki that every employee can edit. An employee adds a page titled "Leave policy — updated" stating that unused leave can be encashed at double pay, written with many keywords from leave questions. For two weeks the assistant quotes it, and 30 employees file encashment requests.
The fix: the assistant now retrieves only from pages in the "HR policy" space, which only the HR team can edit; every answer shows the page owner and last-edited date; and a nightly job flags any edit to policy pages for review. The team also adds 20 "planted document" tests to the eval suite to check that pages outside the approved space are never retrieved.
Follow-up questions to expect
- "How is RAG poisoning different from prompt injection?" — Poisoning plants false facts the model will repeat; injection plants instructions the model will follow. A poisoned document can do both.
- "How much poisoned data is needed?" — Often very little for a targeted attack. Research in 2025 found a roughly fixed, small number of poisoned documents could backdoor models of very different sizes, so "our dataset is huge" is not a defence.
- "Can you detect a poisoned model after training?" — Sometimes, with targeted probes and comparison to a reference model, but not reliably. Prevention through provenance is stronger.