AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How is interpretability different from explainability?


Reading the mechanism versus describing one decisionInterpretability• About the model itself• Built in: scorecards, small trees, GAMs• Faithful — you read how it decides• May cost some accuracyExplainability• About one output• Bolted on: SHAP, LIME, counterfactuals• Approximate — a model of the model• Cheap, and can be confidently wrong
An explainer can point at a proxy the scorecard never uses, so let the interpretable part make the binding decision.

What you need to know

InterpretabilityExplainability
AboutThe model itselfOne decision
WhenBy design or by deep analysisAfter the fact
FaithfulnessHigh — you read the mechanismApproximate; can mislead
CostLimits model choice, or research effortCheap to bolt on
ExamplesLinear models, small trees, scorecards, sparse-autoencoder featuresSHAP, LIME, counterfactuals, citations

Interpretable-by-design models

A credit scorecard gives points for each factor (age of credit history, utilisation, missed payments). Anyone can add up the points and see the decision. Generalised additive models (GAMs) and shallow trees work the same way. You may lose some accuracy compared with a large gradient-boosted model, but on tabular data the gap is often small.

Mechanistic interpretability for LLMs

Research labs try to understand what is computed inside transformers: which attention heads copy information, which directions in activation space represent concepts. Sparse autoencoders split a model's activations into many features that often match human concepts, and changing a feature can change the behaviour. This is an active research area, useful for safety research and debugging. It does not yet give you a per-decision explanation you can show a customer.

Why the difference matters

A post-hoc explanation is a second model of the first one. It can be confidently wrong. If a regulator or a court asks "why was this person rejected?", an interpretable model can answer exactly; an explainer can only give an approximation.

A real-life example

A health insurer builds a claim-triage system. A large gradient-boosted model flags claims that may be fraudulent, with SHAP explanations for investigators. Accuracy is good, but investigators notice the explanations sometimes point to "hospital located in district X" — an uncomfortable proxy for region and income.

The final design: the big model only ranks claims for review order. The binding decision to deny payment uses a 12-factor scorecard reviewed by the compliance team, plus a human investigator. When the team perturbs the district feature, the ranking moves but the scorecard decision cannot, because district is not in it. The explanation given to the customer comes from the scorecard, which is exactly how the decision was made.

Follow-up questions to expect

  • "Is a neural network ever interpretable?" — Parts of one can be studied, but not well enough to give a complete per-decision explanation for a large model today.
  • "When would you accept a black box with explanations?" — When the decision is reversible, reviewed by a person, or low-stakes, and when the accuracy gain is measured and meaningful.
  • "How do you test explanation faithfulness?" — Change the feature the explanation says mattered and check the output actually moves; remove features it says did not matter and check nothing changes.