AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you design safeguards for high-risk use cases (healthcare, finance)?


What you need to know

Why these domains are different

  • The harm from one wrong answer can be physical or financial and irreversible.
  • They are regulated: in the EU AI Act, credit scoring and many medical uses are high-risk; in the US, medical software can be regulated as a device; in India, the RBI and SEBI supervise financial uses (the RBI published a framework for responsible AI in the financial sector in 2025) and medical-device rules can cover diagnostic software.
  • Professionals are accountable for decisions, so the system must support their judgement, not replace it.

The safeguards

  1. Narrow scope — documented tasks and non-goals: "does not diagnose", "does not give personalised investment advice". Out-of-scope requests are refused with a safe route.
  2. Grounded-only generation — from approved clinical guidelines or the bank's own policies, with citations; refuse when evidence does not cover the question.
  3. Deterministic calculations — doses, EMIs, eligibility and risk scores in tested functions, never by the model.
  4. Human decision — a clinician, adviser or credit officer approves before anything reaches a customer or moves money; log who approved what.
  5. Escalation paths — red-flag symptoms, self-harm signals and suspected fraud go to a person immediately.
  6. Change control — pinned versions; every change re-evaluated before release.

Validation

  • Expert-led review of a stratified sample, not just automated metrics.
  • Worst-slice performance: elderly patients, rare conditions, regional languages, thin-file borrowers.
  • Must-pass sets: for example, 100% correct escalation on emergency cases.

A real-life example

A hospital chain builds a symptom-checker for its app. Version 1 is limited to three things: explaining symptoms in plain language from 400 doctor-approved articles, suggesting the right department, and booking an appointment. It never names a diagnosis or a dose.

A deterministic rule list checks every message for red flags (chest pain with sweating, stroke signs, suicidal thoughts, heavy bleeding in pregnancy) before the model runs, and shows an emergency number and the nearest emergency department. On a doctor-written set of 250 emergency cases, the rule list plus the model's own triage escalates 250 of 250. On 1,000 routine cases, 9% are escalated unnecessarily, which the doctors accept as the price of safety. Diagnostic suggestions are planned for version 2, only for doctors, not patients.

Follow-up questions to expect

  • "Why not let the model diagnose if it's accurate on benchmarks?" — Benchmarks don't match your patients, the legal accountability sits with clinicians, and diagnostic software may need regulatory approval as a medical device.
  • "How do you handle a question the knowledge base doesn't cover?" — Refuse clearly and route to a person or a booking, instead of answering from general knowledge.
  • "What would you monitor after launch?" — Escalation rate, clinician override rate, complaints, and a weekly expert review of a random sample.