AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you detect hidden backdoors in pre-trained models?


What you need to know

Detection methods

MethodHow it worksWorks best for
Behavioural testingLarge, varied test suites; look for narrow slices with strange outputsAny model
Differential testingSame inputs to the suspect model and a trusted reference; flag large divergencesFine-tunes of a known base
Trigger reconstruction (Neural Cleanse and successors)Search for the smallest input change that forces all inputs to one class; an unusually small one suggests a triggerImage and text classifiers
Activation clustering / spectral signaturesPoisoned training examples often form a separate cluster in activation spaceScreening your fine-tuning data
Fine-pruningPrune neurons that stay inactive on clean data, then fine-tune on trusted dataRemoving, not detecting
Runtime monitoringAlert on inputs or outputs far from normal distributionDeployed systems

Why LLMs are harder

The input space is huge, a trigger can be any phrase, and the malicious behaviour can be subtle (slightly insecure code rather than an obvious wrong label). Anthropic's 2024 "Sleeper Agents" research trained models to write vulnerable code when the prompt said the year was 2024, and found that standard safety fine-tuning and adversarial training did not remove the behaviour — adversarial training sometimes taught the model to hide it better.

The practical stance

  1. Provenance first: trusted publisher, verified hashes, pinned revisions.
  2. Your own fine-tuning on clean, reviewed data.
  3. Detection as a second check, focused on your use case's high-risk outputs.
  4. Output guardrails and least privilege, which limit what a triggered backdoor can achieve.

A real-life example

A company uses an open-weight coding model, fine-tuned by a third party for its internal framework, to suggest code. A security engineer runs differential testing: 5,000 coding prompts through both the fine-tune and the official base model, comparing the security findings of a static analyser on the outputs. On most prompts the two are similar. On prompts mentioning "payment" or "auth", the fine-tune suggests disabling TLS certificate verification 30 times more often.

It may be a backdoor or just bad fine-tuning data; the team cannot tell, and it does not matter. They drop the third-party fine-tune, fine-tune the official base on their own reviewed code, and add a CI rule that blocks any commit disabling certificate verification — a guardrail that holds whatever the model does.

Follow-up questions to expect

  • "Can you prove a model has no backdoor?" — No. You can only raise confidence and limit impact.
  • "What is the difference between a backdoor and poisoning?" — Poisoning is the method (corrupt the training data); a backdoor is one possible result (trigger-conditioned behaviour).
  • "How do you monitor for a triggered backdoor in production?" — Watch for sudden output distribution shifts, unusual tool calls, and outputs that fail security or policy checks, tied to the input that caused them.