AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How do you build trust in AI-driven products?


What you need to know

The goal is calibrated trust: users rely on the system when it is right and check it when it might be wrong. Too little trust wastes the product; too much trust (automation bias) causes harm.

The levers

  • Show your evidence. Citations and source snippets let users check a claim in seconds.
  • Be calibrated. Say "I'm not sure" or abstain when retrieval is weak. One confident wrong answer costs more trust than several honest "I don't know" replies.
  • Keep the human in control. Draft, then confirm. Preview before send, undo, edit in place.
  • Be consistent. Pin model and prompt versions; run regression evals so quality does not change silently after a vendor update.
  • Be honest about scope. Tell users what the assistant cannot do before they find out through a failure.
  • Disclose AI. Tell people when they are talking to a bot. Several laws, including the EU AI Act's transparency rules, require this.
  • Make feedback cheap. A report button on each answer that reaches a person.

Measuring trust by behaviour

SignalWhat it tells you
Accepted-unchanged rateHow often suggestions are good enough to use as-is
Edit / override rateWhere users disagree with the system
Escalation to humanWhere the bot is not enough
Thumbs-down rateExplicit dissatisfaction
Repeat usageWhether people come back

A slow fall in acceptance rate is often the earliest sign that quality has drifted.

A real-life example

A bank's chatbot answers questions about fees and charges. At launch, answers had no sources, and 38% of chats ended with "talk to an agent". Customers did not believe the bot on money matters.

The team adds a source link under every fee answer ("From: Schedule of Charges, updated 1 July"), makes the bot say "I don't have that information — connecting you to an agent" when retrieval confidence is low, and adds a one-tap "This is wrong" button that files a ticket with the full trace. Within six weeks escalations fall to 21%, and the "This is wrong" reports reveal two stale fee documents, which are fixed within a day.

Follow-up questions to expect

  • "How do you avoid over-trust?" — Show uncertainty, require confirmation on consequential steps, and seed occasional checks. Watch for near-zero override rates, which often mean people have stopped reading.
  • "Does a disclaimer build trust?" — Not much. A small "AI can make mistakes" line doesn't change behaviour. Showing the source and asking for confirmation does.
  • "What happens after a public failure?" — Admit it, fix it, show what changed, and add the case to the eval suite. Hiding it destroys trust faster than the error.