Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is the difference between Human-in-the-Loop and Human-on-the-Loop?
What you need to know
Human-in-the-loop (HITL)
- Person approves before the action happens
- Bounds worst-case harm
- Adds latency and staffing cost
- Degrades into rubber-stamping at volume
Human-on-the-loop (HOTL)
- System acts; person monitors and can intervene
- Scales, low latency
- Catches errors after they happen
- Only as good as monitoring and response time
A third term, human-out-of-the-loop, means full automation with no real-time supervision — acceptable only for low-impact, reversible tasks.
Choosing
| Decision property | Pattern |
|---|---|
| Irreversible or high harm (payments, rejections, medical actions) | In the loop |
| High volume, reversible, well monitored (spam filtering, content ranking) | On the loop |
| Mixed | On the loop for confident cases; in the loop for the low-confidence tail |
The hybrid is usually right: automate the confident, reversible majority with monitoring, and send the uncertain or high-impact minority to a person.
Automation bias
Both patterns fail when reviewers stop really reviewing. Guard against it:
- Measure override rate and time per review. A near-zero override rate and two-second reviews mean the review isn't real.
- Seed known-bad cases into the queue and check reviewers catch them.
- Show evidence, not just the answer, so reviewers can check.
- Limit queue sizes so reviewers have time.
A real-life example
A bank's card fraud system blocks suspicious transactions automatically (on the loop): 20 analysts watch dashboards, sample blocked transactions and can unblock in bulk. Its customer chatbot, which can reverse fees, uses in-the-loop approval for any reversal above Rs 1,000.
An audit finds the fee-reversal reviewers approve 99.6% of requests in an average of 4 seconds. The team seeds 50 fake requests that clearly break policy; reviewers approve 41 of them. The fix: the review screen now shows the policy rule and the customer's reversal history, queues are capped per reviewer per hour, and seeded tests continue monthly. The next month, reviewers catch 46 of 50.
Follow-up questions to expect
- "Where does the EU AI Act stand?" — High-risk systems must be designed so that people can effectively oversee them, including understanding the output, not over-relying on it, and being able to stop the system.
- "How do you reduce in-the-loop cost?" — Use confidence thresholds so only uncertain cases need review, and good tools that show evidence so each review is fast and real.
- "What is the right sampling rate for on-the-loop review?" — Set by risk and by how quickly you need to detect a problem; start higher for new systems and reduce as measured reliability grows.