Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you detect hidden backdoors in pre-trained models?
What you need to know
Detection methods
| Method | How it works | Works best for |
|---|---|---|
| Behavioural testing | Large, varied test suites; look for narrow slices with strange outputs | Any model |
| Differential testing | Same inputs to the suspect model and a trusted reference; flag large divergences | Fine-tunes of a known base |
| Trigger reconstruction (Neural Cleanse and successors) | Search for the smallest input change that forces all inputs to one class; an unusually small one suggests a trigger | Image and text classifiers |
| Activation clustering / spectral signatures | Poisoned training examples often form a separate cluster in activation space | Screening your fine-tuning data |
| Fine-pruning | Prune neurons that stay inactive on clean data, then fine-tune on trusted data | Removing, not detecting |
| Runtime monitoring | Alert on inputs or outputs far from normal distribution | Deployed systems |
Why LLMs are harder
The input space is huge, a trigger can be any phrase, and the malicious behaviour can be subtle (slightly insecure code rather than an obvious wrong label). Anthropic's 2024 "Sleeper Agents" research trained models to write vulnerable code when the prompt said the year was 2024, and found that standard safety fine-tuning and adversarial training did not remove the behaviour — adversarial training sometimes taught the model to hide it better.
The practical stance
- Provenance first: trusted publisher, verified hashes, pinned revisions.
- Your own fine-tuning on clean, reviewed data.
- Detection as a second check, focused on your use case's high-risk outputs.
- Output guardrails and least privilege, which limit what a triggered backdoor can achieve.
A real-life example
A company uses an open-weight coding model, fine-tuned by a third party for its internal framework, to suggest code. A security engineer runs differential testing: 5,000 coding prompts through both the fine-tune and the official base model, comparing the security findings of a static analyser on the outputs. On most prompts the two are similar. On prompts mentioning "payment" or "auth", the fine-tune suggests disabling TLS certificate verification 30 times more often.
It may be a backdoor or just bad fine-tuning data; the team cannot tell, and it does not matter. They drop the third-party fine-tune, fine-tune the official base on their own reviewed code, and add a CI rule that blocks any commit disabling certificate verification — a guardrail that holds whatever the model does.
Follow-up questions to expect
- "Can you prove a model has no backdoor?" — No. You can only raise confidence and limit impact.
- "What is the difference between a backdoor and poisoning?" — Poisoning is the method (corrupt the training data); a backdoor is one possible result (trigger-conditioned behaviour).
- "How do you monitor for a triggered backdoor in production?" — Watch for sudden output distribution shifts, unusual tool calls, and outputs that fail security or policy checks, tied to the input that caused them.