Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is the role of AI governance frameworks in system design?
What you need to know
The main frameworks
| Framework | Type | What it gives you |
|---|---|---|
| NIST AI RMF 1.0 (2023) | Voluntary, US | Four functions: Govern (culture, roles, policies), Map (context and risks of a use case), Measure (test and track risks), Manage (prioritise and respond) |
| NIST AI 600-1, Generative AI Profile (2024) | Voluntary, US | Twelve risks specific to or made worse by generative AI — including confabulation, information security, harmful bias, data privacy, intellectual property, information integrity and value-chain risks — with suggested actions |
| ISO/IEC 42001:2023 | Certifiable standard | An AI management system: policies, roles, risk assessment, controls, internal audit, continual improvement. Works like ISO 27001 for security. |
| EU AI Act | Law | Risk tiers with binding obligations, especially for high-risk systems and general-purpose models |
| OWASP Top 10 for LLM Applications | Technical guidance | The security risks your controls must address |
Where they touch design
- Intake: every use case gets a risk tier. The tier decides mandatory controls.
- Data: provenance and lawful basis recorded before training or indexing.
- Evaluation: required evals, fairness slices and red-team results, with thresholds.
- Documentation: model and system cards with intended use, limits and results.
- Oversight: a designed human checkpoint for higher tiers.
- Operations: logging, monitoring, incident response, rollback.
- Lifecycle: re-review dates and criteria for switching off.
Wired in, not run beside
The failure mode is governance as a separate committee producing documents about systems that already shipped. The fix:
- The risk tier is a field in the service catalogue.
- The release checklist is a CI gate.
- The model card is generated from the latest eval run.
- The accountable owner can block a launch.
A real-life example
A mid-size bank adopts NIST AI RMF as its internal structure and aims for ISO/IEC 42001 certification. For its new customer chatbot, Map produces a one-page risk profile: customers may act on wrong fee information; personal data flows to a model vendor; the bot could be manipulated to reveal account data. Measure turns these into three eval suites (fee accuracy, PII leakage, injection resistance) with thresholds. Manage assigns owners and a monthly review. Govern names the head of digital banking as accountable owner.
When a vendor model upgrade drops fee accuracy from 97% to 91%, the CI gate blocks the release automatically. Nobody has to remember to check, which is the point.
Follow-up questions to expect
- "Is NIST AI RMF mandatory?" — No, it is voluntary guidance, but it is widely used as a common language and is referenced in contracts and procurement.
- "What does ISO 42001 certification actually prove?" — That you run a management system for AI risk with audits and improvement; it does not certify that any specific model is safe.
- "How do you keep governance lightweight?" — Tier by risk: internal low-risk tools get a short checklist; only high-risk systems get full review.