Course Content
AI Safety & Guardrails
5 sections · 50 lessons
When should human oversight be required in AI workflows?
What you need to know
Triggers for oversight
- Irreversible or expensive actions: payments, deletions, external emails, production changes, publishing.
- Significant effects on a person: hiring, credit, insurance, benefits, admissions, account suspension. GDPR and the EU AI Act's high-risk rules apply here.
- Regulated advice: medical, legal, financial.
- Safety-critical steps.
- Low confidence: weak retrieval, low classifier confidence, disagreement between samples.
- A guardrail fired or the input looks like manipulation.
- New deployment: higher review rates at first, lowered as reliability is measured.
Where oversight does not pay
Mandatory review on high-volume, reversible, low-harm actions creates queues, delay and reviewers who approve without reading. That weakens oversight on the decisions that matter. Treat reviewer attention as a limited budget.
A policy enforced in code
1from dataclasses import dataclass23@dataclass4class Action:5 name: str6 irreversible: bool = False7 external: bool = False8 amount: float = 09 affects_person: bool = False1011AUTO_LIMIT = {"low": 2000, "medium": 500, "high": 0}1213def oversight(action, risk_tier, confidence, guardrail_fired=False):14 if action.affects_person or action.irreversible:15 return "approve"16 if action.external or guardrail_fired or confidence < 0.7:17 return "approve"18 if action.amount > AUTO_LIMIT[risk_tier]:19 return "approve"20 return "auto" # still sampled for audit2122print(oversight(Action("label_email"), "low", 0.93)) # auto23print(oversight(Action("send_reply", external=True), "low", 0.97)) # approve24print(oversight(Action("refund", amount=1500), "medium", 0.9)) # approve25print(oversight(Action("reject_candidate", affects_person=True), "high", 0.99)) # approveThe tool layer calls oversight() before executing any action. The model cannot skip it by phrasing. Even a 99%-confident candidate rejection goes to a person, because it affects someone's job prospects. The "auto" path is still sampled so you learn when thresholds are wrong.
A real-life example
An email assistant for a law firm's clients can label, summarise, draft and send email. The first design asked the user to approve every action, including labels — about 300 approvals a day per user. Within two weeks users were clicking "approve all" without reading.
The redesign: labelling and summarising run automatically (reversible, internal); drafting is automatic but never sends; sending always needs the user's click with the recipient list shown, and any email to a new external domain or with attachments shows an extra warning; anything triggered while processing an email from an unknown sender needs approval, because that content is untrusted. Approvals fall to about 15 a day, and users now read them.
Follow-up questions to expect
- "How do you set the confidence threshold?" — On a labelled set: pick the threshold where auto-approved cases meet your error target, and check it again after model changes.
- "What if reviewers are the bottleneck?" — Reduce what needs review with better thresholds and scope, improve reviewer tools, and never solve it by quietly removing review from high-impact actions.
- "Should oversight ever decrease?" — Yes, based on measured reliability over time, with sampling kept on the automatic path.