AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

When should human oversight be required in AI workflows?


The oversight policy the tool layer runsIrreversible oraffects a person?Yes: human approvesNo: external,unsure or flagged?Yes: human approvesNo: over theamount limit?Yes: human approvesNo: auto,sampled for audit
Reviewer attention is a budget, so spend it on the irreversible and uncertain, and sample the rest.

What you need to know

Triggers for oversight

  • Irreversible or expensive actions: payments, deletions, external emails, production changes, publishing.
  • Significant effects on a person: hiring, credit, insurance, benefits, admissions, account suspension. GDPR and the EU AI Act's high-risk rules apply here.
  • Regulated advice: medical, legal, financial.
  • Safety-critical steps.
  • Low confidence: weak retrieval, low classifier confidence, disagreement between samples.
  • A guardrail fired or the input looks like manipulation.
  • New deployment: higher review rates at first, lowered as reliability is measured.

Where oversight does not pay

Mandatory review on high-volume, reversible, low-harm actions creates queues, delay and reviewers who approve without reading. That weakens oversight on the decisions that matter. Treat reviewer attention as a limited budget.

A policy enforced in code

Python
from dataclasses import dataclass@dataclassclass Action:    name: str    irreversible: bool = False    external: bool = False    amount: float = 0    affects_person: bool = FalseAUTO_LIMIT = {"low": 2000, "medium": 500, "high": 0}def oversight(action, risk_tier, confidence, guardrail_fired=False):    if action.affects_person or action.irreversible:        return "approve"    if action.external or guardrail_fired or confidence < 0.7:        return "approve"    if action.amount > AUTO_LIMIT[risk_tier]:        return "approve"    return "auto"            # still sampled for auditprint(oversight(Action("label_email"), "low", 0.93))                  # autoprint(oversight(Action("send_reply", external=True), "low", 0.97))    # approveprint(oversight(Action("refund", amount=1500), "medium", 0.9))        # approveprint(oversight(Action("reject_candidate", affects_person=True), "high", 0.99))  # approve

The tool layer calls oversight() before executing any action. The model cannot skip it by phrasing. Even a 99%-confident candidate rejection goes to a person, because it affects someone's job prospects. The "auto" path is still sampled so you learn when thresholds are wrong.

A real-life example

An email assistant for a law firm's clients can label, summarise, draft and send email. The first design asked the user to approve every action, including labels — about 300 approvals a day per user. Within two weeks users were clicking "approve all" without reading.

The redesign: labelling and summarising run automatically (reversible, internal); drafting is automatic but never sends; sending always needs the user's click with the recipient list shown, and any email to a new external domain or with attachments shows an extra warning; anything triggered while processing an email from an unknown sender needs approval, because that content is untrusted. Approvals fall to about 15 a day, and users now read them.

Follow-up questions to expect

  • "How do you set the confidence threshold?" — On a labelled set: pick the threshold where auto-approved cases meet your error target, and check it again after model changes.
  • "What if reviewers are the bottleneck?" — Reduce what needs review with better thresholds and scope, improve reviewer tools, and never solve it by quietly removing review from high-impact actions.
  • "Should oversight ever decrease?" — Yes, based on measured reliability over time, with sampling kept on the automatic path.