Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is responsible AI, and how do organizations implement it?
What you need to know
Most companies publish similar principles. The difference between them is whether the principles change engineering decisions.
The lifecycle controls
- Intake and risk tiering — every use case gets a tier (internal productivity tool vs. a decision about a person's loan or job). The tier decides which controls are mandatory.
- Design review — data sources and lawful basis, the harm model (who could be hurt and how), the human-oversight point.
- Documentation — a model card or system card: intended use, out-of-scope use, eval results, known limitations.
- Pre-launch gates — eval thresholds, red-team results, a fairness slice report, a privacy review, and sign-off by a named accountable owner.
- Runtime controls — guardrails, human review where required, monitoring, audit logs.
- After launch — incident response, user appeals, periodic re-review, and criteria for switching it off.
The frameworks, briefly
- NIST AI RMF 1.0 (US, voluntary): four functions — Govern, Map, Measure, Manage. NIST's Generative AI Profile (NIST AI 600-1, 2024) lists risks specific to generative AI, such as confabulation, information security and harmful bias, with suggested actions.
- ISO/IEC 42001:2023: a certifiable AI management system standard, like ISO 27001 for security. Enterprise buyers increasingly ask for it.
- EU AI Act: law, not guidance. It sorts systems by risk tier and puts binding duties on high-risk ones.
- India: no single AI law yet; the DPDP Act 2023 governs personal data, and MeitY has issued AI governance guidelines and advisories.
What "enforced" looks like
- The risk tier is a field in the service catalogue.
- The release checklist is a CI gate that fails the build.
- The model card is generated from the latest eval run, not typed by hand.
- The accountable owner has budget and authority to stop a launch.
A real-life example
A lending app wants an LLM to summarise bank statements for loan officers. At intake it is tiered "high" because it influences credit decisions. That tier requires: a fairness report across gender, age band and state; a red-team of 150 cases including injected text in uploaded PDFs; officers must see the source rows for every figure; and the head of credit risk signs off.
The fairness report finds summaries for statements in Hindi and Tamil drop key salary credits 9% of the time versus 2% for English. Launch is delayed three weeks to fix extraction. Without the gate, this would have surfaced as a pattern of rejected applicants from two states.
Follow-up questions to expect
- "Who owns responsible AI in a company?" — A central team sets the policy and tooling; each product team owns its system's compliance, with a named accountable owner per system.
- "How do you keep it from slowing everything down?" — Tiering. Low-risk internal tools get a light checklist; only high-risk systems get the full review.
- "What is a model card?" — A short document that states what a model is for, how it was evaluated, on which groups, and its known limitations and out-of-scope uses.