AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What is the role of AI governance frameworks in system design?


NIST AI RMF applied to a bank chatbotGovern — named owner, policies, cultureMap — who can be harmed, and howMeasure — eval suites with thresholdsManage — owners, reviews, response
When the vendor upgrade dropped fee accuracy to 91%, the CI gate blocked it without anyone remembering to check.

What you need to know

The main frameworks

FrameworkTypeWhat it gives you
NIST AI RMF 1.0 (2023)Voluntary, USFour functions: Govern (culture, roles, policies), Map (context and risks of a use case), Measure (test and track risks), Manage (prioritise and respond)
NIST AI 600-1, Generative AI Profile (2024)Voluntary, USTwelve risks specific to or made worse by generative AI — including confabulation, information security, harmful bias, data privacy, intellectual property, information integrity and value-chain risks — with suggested actions
ISO/IEC 42001:2023Certifiable standardAn AI management system: policies, roles, risk assessment, controls, internal audit, continual improvement. Works like ISO 27001 for security.
EU AI ActLawRisk tiers with binding obligations, especially for high-risk systems and general-purpose models
OWASP Top 10 for LLM ApplicationsTechnical guidanceThe security risks your controls must address

Where they touch design

  • Intake: every use case gets a risk tier. The tier decides mandatory controls.
  • Data: provenance and lawful basis recorded before training or indexing.
  • Evaluation: required evals, fairness slices and red-team results, with thresholds.
  • Documentation: model and system cards with intended use, limits and results.
  • Oversight: a designed human checkpoint for higher tiers.
  • Operations: logging, monitoring, incident response, rollback.
  • Lifecycle: re-review dates and criteria for switching off.

Wired in, not run beside

The failure mode is governance as a separate committee producing documents about systems that already shipped. The fix:

  • The risk tier is a field in the service catalogue.
  • The release checklist is a CI gate.
  • The model card is generated from the latest eval run.
  • The accountable owner can block a launch.

A real-life example

A mid-size bank adopts NIST AI RMF as its internal structure and aims for ISO/IEC 42001 certification. For its new customer chatbot, Map produces a one-page risk profile: customers may act on wrong fee information; personal data flows to a model vendor; the bot could be manipulated to reveal account data. Measure turns these into three eval suites (fee accuracy, PII leakage, injection resistance) with thresholds. Manage assigns owners and a monthly review. Govern names the head of digital banking as accountable owner.

When a vendor model upgrade drops fee accuracy from 97% to 91%, the CI gate blocks the release automatically. Nobody has to remember to check, which is the point.

Follow-up questions to expect

  • "Is NIST AI RMF mandatory?" — No, it is voluntary guidance, but it is widely used as a common language and is referenced in contracts and procurement.
  • "What does ISO 42001 certification actually prove?" — That you run a management system for AI risk with audits and improvement; it does not certify that any specific model is safe.
  • "How do you keep governance lightweight?" — Tier by risk: internal low-risk tools get a short checklist; only high-risk systems get full review.