Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What is explainability, and why do users and regulators care about it?
What you need to know
Techniques for classical ML
- Feature attribution: SHAP assigns each feature a contribution to one prediction; LIME fits a simple local model around the prediction.
- Counterfactual explanations: the smallest change to the input that flips the decision. People find these the most actionable.
- Global surrogates: a small decision tree trained to imitate the big model, for an overview.
Explanations for LLM systems
- Citations that link each claim to a retrieved passage.
- Trace: which documents were retrieved, which tools ran with which arguments.
- Policy reason: "Declined because the request asked for another customer's data."
- Model-stated reasoning: useful for debugging, but it is generated text, not a log of the computation.
Who cares, and why
- Users: they cannot trust an answer they cannot check. A cited source turns a plausible answer into a usable one.
- Operators and support: an unexplained decision cannot be debugged or appealed.
- Regulators: GDPR gives people rights to meaningful information about the logic of solely automated decisions with significant effects. The EU AI Act requires high-risk systems to be transparent enough for deployers to interpret outputs and use them correctly. Lending and insurance rules in many countries require reasons for adverse decisions.
A real-life example
A digital lender rejects a personal-loan application. The old message said "Your application does not meet our criteria." Complaints were high and the support team could not answer them.
The new system gives the top reasons from the credit model's SHAP values in plain language — "High existing EMIs relative to income (48%)" and "Two missed credit-card payments in the last 6 months" — plus a counterfactual: "Applications with EMIs below 40% of income are usually approved." The team checks faithfulness: for 500 rejected cases they reduce the EMI ratio in the input and confirm the decision actually flips in 91% of cases where the explanation named it. Complaint tickets per rejection fall by about a third.
Follow-up questions to expect
- "Is chain-of-thought an explanation?" — It is a useful hint, not a faithful record. Prefer explanations you can verify, such as a citation that either supports the claim or doesn't.
- "What is the risk of giving explanations?" — Explanations can help people game the system, and can leak model or data details. Give reasons that are true and actionable, not the full model.
- "How do you test an explanation?" — Perturb the feature it names and check the output really changes; check that a cited passage actually entails the sentence.