Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What are the risks of using third-party or open-source models?
What you need to know
OWASP lists this as LLM03: Supply Chain. The chain includes base models, fine-tuned variants, LoRA adapters, tokenisers, datasets, Python packages and model-serving code.
Risks from open-weight models
- Unsafe file formats: PyTorch's classic
.bin/.ptfiles use Python pickle, which can execute arbitrary code when loaded. safetensors stores only tensors and cannot run code. Recent PyTorch defaultstorch.loadtoweights_only=True, which helps, but do not rely on it for untrusted files. - Malicious or tampered models: security researchers have found models on public hubs that open a reverse shell on load, and typosquatted names that imitate popular publishers.
- Backdoors: behaviour that passes benchmarks and misfires on a trigger.
- Licences: "open" can mean non-commercial only, an acceptable-use policy, or conditions tied to company size or attribution. Read the licence file, not the README badge.
- Training data unknowns: copyrighted or personal data that can surface in outputs.
- Contaminated benchmarks: the model card's scores may reflect test data leaked into training.
- Maintenance: no security fixes; a repository can disappear.
Risks from hosted third-party APIs
- Data retention and whether your data is used for training.
- Processing region and cross-border transfer.
- Silent updates: a model alias changes behaviour and breaks your evals.
- Deprecation schedules and availability.
Controls
- Source — pull only from verified publishers; record the exact revision (commit hash).
- Verify — check file hashes and signatures where available; mirror approved artifacts into an internal registry.
- Scan and isolate — prefer safetensors; scan with a model-file scanner; load untrusted files in a sandbox with no network or secrets.
- Evaluate — run your own task evals and red-team suite; don't trust the card.
- Review — legal reads the licence; security records the model in an AI bill of materials.
- Abstract — keep a provider interface so you can switch models when terms or quality change.
A real-life example
A fintech team picks a fine-tuned 8B model from a public hub for classifying customer complaints. It scores 94% on the card's benchmark. The platform team's checks find: the weights are pickle files from an individual uploader with 30 followers; the licence of the base model forbids use for "credit decisions", which one downstream workflow does; and on the team's own 1,000 real complaints, accuracy is 81%.
They switch to the official base model in safetensors from the original publisher, fine-tune it themselves on their labelled data with a pinned revision, reach 90%, and store the model with its hash and licence in the internal registry. The whole detour cost a week; a pickle exploit or a licence breach would have cost far more.
Follow-up questions to expect
- "Is a model from a big company automatically safe?" — Safer to load, but you still need to check the licence, your own evals, and the exact revision you pin.
- "What is an AI BOM?" — An AI bill of materials: a record of every model, dataset, adapter and library in a system, with versions, sources and licences, so you can respond when one of them has a problem.
- "How do you handle a vendor model update?" — Pin a dated model version, not a floating alias, and run the eval suite before switching.