Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do you build trust in AI-driven products?
What you need to know
The goal is calibrated trust: users rely on the system when it is right and check it when it might be wrong. Too little trust wastes the product; too much trust (automation bias) causes harm.
The levers
- Show your evidence. Citations and source snippets let users check a claim in seconds.
- Be calibrated. Say "I'm not sure" or abstain when retrieval is weak. One confident wrong answer costs more trust than several honest "I don't know" replies.
- Keep the human in control. Draft, then confirm. Preview before send, undo, edit in place.
- Be consistent. Pin model and prompt versions; run regression evals so quality does not change silently after a vendor update.
- Be honest about scope. Tell users what the assistant cannot do before they find out through a failure.
- Disclose AI. Tell people when they are talking to a bot. Several laws, including the EU AI Act's transparency rules, require this.
- Make feedback cheap. A report button on each answer that reaches a person.
Measuring trust by behaviour
| Signal | What it tells you |
|---|---|
| Accepted-unchanged rate | How often suggestions are good enough to use as-is |
| Edit / override rate | Where users disagree with the system |
| Escalation to human | Where the bot is not enough |
| Thumbs-down rate | Explicit dissatisfaction |
| Repeat usage | Whether people come back |
A slow fall in acceptance rate is often the earliest sign that quality has drifted.
A real-life example
A bank's chatbot answers questions about fees and charges. At launch, answers had no sources, and 38% of chats ended with "talk to an agent". Customers did not believe the bot on money matters.
The team adds a source link under every fee answer ("From: Schedule of Charges, updated 1 July"), makes the bot say "I don't have that information — connecting you to an agent" when retrieval confidence is low, and adds a one-tap "This is wrong" button that files a ticket with the full trace. Within six weeks escalations fall to 21%, and the "This is wrong" reports reveal two stale fee documents, which are fixed within a day.
Follow-up questions to expect
- "How do you avoid over-trust?" — Show uncertainty, require confirmation on consequential steps, and seed occasional checks. Watch for near-zero override rates, which often mean people have stopped reading.
- "Does a disclaimer build trust?" — Not much. A small "AI can make mistakes" line doesn't change behaviour. Showing the source and asking for confirmation does.
- "What happens after a public failure?" — Admit it, fix it, show what changed, and add the case to the eval suite. Hiding it destroys trust faster than the error.