Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How is interpretability different from explainability?
What you need to know
| Interpretability | Explainability | |
|---|---|---|
| About | The model itself | One decision |
| When | By design or by deep analysis | After the fact |
| Faithfulness | High — you read the mechanism | Approximate; can mislead |
| Cost | Limits model choice, or research effort | Cheap to bolt on |
| Examples | Linear models, small trees, scorecards, sparse-autoencoder features | SHAP, LIME, counterfactuals, citations |
Interpretable-by-design models
A credit scorecard gives points for each factor (age of credit history, utilisation, missed payments). Anyone can add up the points and see the decision. Generalised additive models (GAMs) and shallow trees work the same way. You may lose some accuracy compared with a large gradient-boosted model, but on tabular data the gap is often small.
Mechanistic interpretability for LLMs
Research labs try to understand what is computed inside transformers: which attention heads copy information, which directions in activation space represent concepts. Sparse autoencoders split a model's activations into many features that often match human concepts, and changing a feature can change the behaviour. This is an active research area, useful for safety research and debugging. It does not yet give you a per-decision explanation you can show a customer.
Why the difference matters
A post-hoc explanation is a second model of the first one. It can be confidently wrong. If a regulator or a court asks "why was this person rejected?", an interpretable model can answer exactly; an explainer can only give an approximation.
A real-life example
A health insurer builds a claim-triage system. A large gradient-boosted model flags claims that may be fraudulent, with SHAP explanations for investigators. Accuracy is good, but investigators notice the explanations sometimes point to "hospital located in district X" — an uncomfortable proxy for region and income.
The final design: the big model only ranks claims for review order. The binding decision to deny payment uses a 12-factor scorecard reviewed by the compliance team, plus a human investigator. When the team perturbs the district feature, the ranking moves but the scorecard decision cannot, because district is not in it. The explanation given to the customer comes from the scorecard, which is exactly how the decision was made.
Follow-up questions to expect
- "Is a neural network ever interpretable?" — Parts of one can be studied, but not well enough to give a complete per-decision explanation for a large model today.
- "When would you accept a black box with explanations?" — When the decision is reversible, reviewed by a person, or low-stakes, and when the accuracy gain is measured and meaningful.
- "How do you test explanation faithfulness?" — Change the feature the explanation says mattered and check the output actually moves; remove features it says did not matter and check nothing changes.