AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

What is the difference between Human-in-the-Loop and Human-on-the-Loop?


Approving first versus watching and stepping inHuman-in-the-loop• Nothing happens until a person approves• Bounds the worst case• Slow and costly at volume• Drifts into rubber-stampingHuman-on-the-loop• System acts, a person supervises• Scales with low latency• Catches errors after the fact• Only as good as the monitoring
Reviewers approved 41 of 50 seeded bad requests, so measure override rates before claiming either pattern works.

What you need to know

Human-in-the-loop (HITL)

  • Person approves before the action happens
  • Bounds worst-case harm
  • Adds latency and staffing cost
  • Degrades into rubber-stamping at volume

Human-on-the-loop (HOTL)

  • System acts; person monitors and can intervene
  • Scales, low latency
  • Catches errors after they happen
  • Only as good as monitoring and response time

A third term, human-out-of-the-loop, means full automation with no real-time supervision — acceptable only for low-impact, reversible tasks.

Choosing

Decision propertyPattern
Irreversible or high harm (payments, rejections, medical actions)In the loop
High volume, reversible, well monitored (spam filtering, content ranking)On the loop
MixedOn the loop for confident cases; in the loop for the low-confidence tail

The hybrid is usually right: automate the confident, reversible majority with monitoring, and send the uncertain or high-impact minority to a person.

Automation bias

Both patterns fail when reviewers stop really reviewing. Guard against it:

  • Measure override rate and time per review. A near-zero override rate and two-second reviews mean the review isn't real.
  • Seed known-bad cases into the queue and check reviewers catch them.
  • Show evidence, not just the answer, so reviewers can check.
  • Limit queue sizes so reviewers have time.

A real-life example

A bank's card fraud system blocks suspicious transactions automatically (on the loop): 20 analysts watch dashboards, sample blocked transactions and can unblock in bulk. Its customer chatbot, which can reverse fees, uses in-the-loop approval for any reversal above Rs 1,000.

An audit finds the fee-reversal reviewers approve 99.6% of requests in an average of 4 seconds. The team seeds 50 fake requests that clearly break policy; reviewers approve 41 of them. The fix: the review screen now shows the policy rule and the customer's reversal history, queues are capped per reviewer per hour, and seeded tests continue monthly. The next month, reviewers catch 46 of 50.

Follow-up questions to expect

  • "Where does the EU AI Act stand?" — High-risk systems must be designed so that people can effectively oversee them, including understanding the output, not over-relying on it, and being able to stop the system.
  • "How do you reduce in-the-loop cost?" — Use confidence thresholds so only uncertain cases need review, and good tools that show evidence so each review is fast and real.
  • "What is the right sampling rate for on-the-loop review?" — Set by risk and by how quickly you need to detect a problem; start higher for new systems and reduce as measured reliability grows.