AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What are the risks of using third-party or open-source models?


What you need to know

OWASP lists this as LLM03: Supply Chain. The chain includes base models, fine-tuned variants, LoRA adapters, tokenisers, datasets, Python packages and model-serving code.

Risks from open-weight models

  • Unsafe file formats: PyTorch's classic .bin/.pt files use Python pickle, which can execute arbitrary code when loaded. safetensors stores only tensors and cannot run code. Recent PyTorch defaults torch.load to weights_only=True, which helps, but do not rely on it for untrusted files.
  • Malicious or tampered models: security researchers have found models on public hubs that open a reverse shell on load, and typosquatted names that imitate popular publishers.
  • Backdoors: behaviour that passes benchmarks and misfires on a trigger.
  • Licences: "open" can mean non-commercial only, an acceptable-use policy, or conditions tied to company size or attribution. Read the licence file, not the README badge.
  • Training data unknowns: copyrighted or personal data that can surface in outputs.
  • Contaminated benchmarks: the model card's scores may reflect test data leaked into training.
  • Maintenance: no security fixes; a repository can disappear.

Risks from hosted third-party APIs

  • Data retention and whether your data is used for training.
  • Processing region and cross-border transfer.
  • Silent updates: a model alias changes behaviour and breaks your evals.
  • Deprecation schedules and availability.

Controls

  1. Source — pull only from verified publishers; record the exact revision (commit hash).
  2. Verify — check file hashes and signatures where available; mirror approved artifacts into an internal registry.
  3. Scan and isolate — prefer safetensors; scan with a model-file scanner; load untrusted files in a sandbox with no network or secrets.
  4. Evaluate — run your own task evals and red-team suite; don't trust the card.
  5. Review — legal reads the licence; security records the model in an AI bill of materials.
  6. Abstract — keep a provider interface so you can switch models when terms or quality change.

A real-life example

A fintech team picks a fine-tuned 8B model from a public hub for classifying customer complaints. It scores 94% on the card's benchmark. The platform team's checks find: the weights are pickle files from an individual uploader with 30 followers; the licence of the base model forbids use for "credit decisions", which one downstream workflow does; and on the team's own 1,000 real complaints, accuracy is 81%.

They switch to the official base model in safetensors from the original publisher, fine-tune it themselves on their labelled data with a pinned revision, reach 90%, and store the model with its hash and licence in the internal registry. The whole detour cost a week; a pickle exploit or a licence breach would have cost far more.

Follow-up questions to expect

  • "Is a model from a big company automatically safe?" — Safer to load, but you still need to check the licence, your own evals, and the exact revision you pin.
  • "What is an AI BOM?" — An AI bill of materials: a record of every model, dataset, adapter and library in a system, with versions, sources and licences, so you can respond when one of them has a problem.
  • "How do you handle a vendor model update?" — Pin a dated model version, not a floating alias, and run the eval suite before switching.