Course Content
AI Ethics and Governance
3 sections · 7 lessons
Explainability Tools (LIME, SHAP, InterpretML)
Ribeiro, Singh and Guestrin trained a classifier to tell huskies from wolves and deliberately rigged the training set: nearly every wolf photo had snow in the background, nearly every husky photo did not. The classifier reached high accuracy on a held-out set drawn the same way. They then showed it to 27 graduate students who had taken at least one machine learning course. Before seeing any explanation, 10 of them said they trusted the model. Then the researchers highlighted the image regions the classifier was actually using. The highlighted region was the snow. Afterwards only 3 still trusted it, and 25 of the 27 named the background as the feature.
Accuracy told them nothing. The held-out set had the same snow artefact as the training set, so the model scored well by exploiting it. The only thing that exposed the failure was a picture of which pixels the decision depended on.
That is the entire case for explainability. It is not a compliance checkbox and it is not about making users comfortable. It is a debugging instrument that catches the specific class of failure that validation metrics are structurally blind to: the model is right for the wrong reason, and your test set shares the reason.
What "explanation" actually means — three axes
People use the word for at least six different artefacts. Sort them along three axes and the confusion drops away.
| Axis | Option A | Option B | Why it matters |
|---|---|---|---|
| Scope | Global — how the model behaves overall | Local — why this prediction, for this person | A regulator asking "what does your model do" needs global. A rejected applicant needs local. |
| Timing | Intrinsic — the model is readable by construction | Post-hoc — a second procedure explains a fitted black box | Post-hoc explanations approximate; intrinsic ones are the truth |
| Coupling | Model-specific — exploits internal structure | Model-agnostic — only needs a predict function | Model-specific is far faster and often exact; agnostic works on anything, including a vendor API |
The crucial distinction hides in the second row. A post-hoc explanation is a model of your model. It has its own error. When LIME reports that income contributed +0.31, that is LIME's estimate of what your model did, not a readout from inside your model. Treat it as a measurement with uncertainty, not as ground truth.
A post-hoc explanation is a second model fitted to the first one. Its output can be wrong in exactly the way any model can be wrong.
LIME: build a simple model that is right near one point
The intuition
A gradient-boosted forest with 400 trees is hopeless to describe globally. But zoom in far enough around a single applicant and almost any decision surface looks close to flat. LIME (Local Interpretable Model-agnostic Explanations) exploits that: it fits a small sparse linear model that agrees with the black box in the neighbourhood of one instance, then reports the linear model's coefficients as the explanation.
The procedure for one applicant:
- Take the instance you want to explain.
- Generate perturbed copies — say 5,000 of them — by sampling around it. For tabular data, resample each feature from the training distribution; for text, drop random words; for images, switch superpixels on and off.
- Ask the black box for a prediction on every perturbed copy.
- Weight each copy by how close it is to the original, using an exponential kernel. Copies near the instance dominate the fit; distant ones barely count.
- Fit a weighted sparse linear regression (usually via LASSO, keeping the top k features) from perturbed inputs to black-box outputs.
- Report those k coefficients.
black-box decision boundary . . ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ . . ~ o o . [X] ~ o o o . . ~ o o . ~ o o ------ <-- LIME's straight line: wrong globally, close enough right next to [X]Running it
1from lime.lime_tabular import LimeTabularExplainer23explainer = LimeTabularExplainer(4 training_data=X_train.values,5 feature_names=list(X_train.columns),6 class_names=["reject", "approve"],7 discretize_continuous=True, # bins numerics; makes rules readable8 random_state=0, # without this, reruns differ9)1011exp = explainer.explain_instance(12 data_row=X_test.iloc[42].values,13 predict_fn=model.predict_proba,14 num_features=6,15 num_samples=5000, # more samples = more stable, linearly more compute16)17for feature, weight in exp.as_list():18 print(f"{feature:<34} {weight:+.3f}")A typical output for a rejected applicant:
debt_to_income > 0.42 -0.184months_at_address <= 11 -0.096num_recent_enquiries > 3 -0.071income > 52000 +0.058account_age_years > 4 +0.041employment = permanent +0.022Read that as: within LIME's local linear approximation, the debt-to-income ratio being above 0.42 pushed the approval probability down by about 0.18 relative to the local baseline. It is a statement about the neighbourhood, not a global claim about the model.
Where LIME breaks
Instability. The perturbations are random, so two runs on the same instance give different numbers, sometimes with a different feature at the top. Alvarez-Melis and Jaakkola (2018) documented that two nearly identical inputs can also receive noticeably different explanations. Always fix the seed, always raise num_samples until the explanation stops moving, and check by running it five times and comparing the top-3 sets.
The neighbourhood is a hyperparameter. The kernel width decides what "local" means. Too narrow and you fit noise; too wide and you are approximating a curved surface with a plane and the coefficients become meaningless. There is no principled default, and results are genuinely sensitive to it.
Unrealistic perturbations. Sampling features independently produces rows that cannot exist — age 22 with 30 years of employment history. The black box is then queried far off the data manifold, where its behaviour is arbitrary, and that arbitrary behaviour becomes part of your explanation.
SHAP: split the credit fairly, with a guarantee
The idea, from game theory
Treat the features as players in a cooperative game. The "payout" is the prediction minus the average prediction. How much of that payout does each feature deserve? Lloyd Shapley answered this in 1953 for economics, and there is exactly one allocation satisfying a short list of fairness axioms — efficiency (the parts sum to the whole), symmetry (features that always contribute identically get identical credit), dummy (a feature that never changes anything gets zero) and additivity.
In words: average the feature's marginal contribution over every possible order in which the features could be revealed.
Computing one by hand
Three features — Income, Age, Debt — and a loan model. Let v(S) be the model's expected output when only the features in S are known. Suppose:
| Known features S | v(S) | Known features S | v(S) |
|---|---|---|---|
| {} (base rate) | 0.30 | {Income, Age} | 0.65 |
| {Income} | 0.60 | {Income, Debt} | 0.50 |
| {Age} | 0.35 | {Age, Debt} | 0.22 |
| {Debt} | 0.20 | {Income, Age, Debt} | 0.55 |
With three features there are 3!=6 orders. Walk each one and record what each feature added when it arrived.
| Order | Income adds | Age adds | Debt adds |
|---|---|---|---|
| I → A → D | 0.60 − 0.30 = +0.30 | 0.65 − 0.60 = +0.05 | 0.55 − 0.65 = −0.10 |
| I → D → A | +0.30 | 0.55 − 0.50 = +0.05 | 0.50 − 0.60 = −0.10 |
| A → I → D | 0.65 − 0.35 = +0.30 | 0.35 − 0.30 = +0.05 | 0.55 − 0.65 = −0.10 |
| A → D → I | 0.55 − 0.22 = +0.33 | +0.05 | 0.22 − 0.35 = −0.13 |
| D → I → A | 0.50 − 0.20 = +0.30 | 0.55 − 0.50 = +0.05 | 0.20 − 0.30 = −0.10 |
| D → A → I | 0.55 − 0.22 = +0.33 | 0.22 − 0.20 = +0.02 | −0.10 |
Average each column:
- ϕIncome=(0.30+0.30+0.30+0.33+0.30+0.33)/6=1.86/6=+0.310
- ϕAge=(0.05+0.05+0.05+0.05+0.05+0.02)/6=0.27/6=+0.045
- ϕDebt=(−0.10−0.10−0.10−0.13−0.10−0.10)/6=−0.63/6=−0.105
Check the efficiency axiom: 0.310+0.045−0.105=0.250, and the prediction minus the base rate is 0.55−0.30=0.250. They match exactly, and they will always match. That guarantee is what LIME does not give you — LIME's coefficients do not have to add up to anything in particular.
Notice also that order matters to the individual terms even though it does not matter to the average. Debt contributes −0.10 in most orders but −0.13 when it arrives after Age alone. That gap is an interaction: Debt hurts slightly more when Age is already known. Averaging over orders is precisely how Shapley handles interactions rather than ignoring them.
Running it
Exact computation needs 2n subsets, which is impossible past about 20 features. Two practical routes exist.
1import shap23# Route 1: TreeSHAP - exact, polynomial time, for tree ensembles only4explainer = shap.TreeExplainer(gbm_model)5sv = explainer(X_test) # thousands of rows in seconds67shap.plots.waterfall(sv[42]) # local: one prediction, decomposed8shap.plots.beeswarm(sv) # global: every row, every feature9shap.plots.scatter(sv[:, "debt_to_income"]) # the learned shape of one feature1011# Route 2: KernelSHAP - model-agnostic, approximate, slow12kernel = shap.KernelExplainer(any_model.predict_proba,13 shap.sample(X_train, 100)) # background set14sv_k = kernel.shap_values(X_test.iloc[:50], nsamples=1000)TreeSHAP (Lundberg et al., 2020) is the reason SHAP took over industry practice: it computes exact Shapley values for tree ensembles in time polynomial in tree depth rather than exponential in features, which turned a research idea into something you can run on a production dataset.
Where SHAP breaks
The background set is a modelling choice, not a detail. "The model's output when a feature is unknown" is undefined; SHAP approximates it by marginalising over a background dataset. Change the background from "all applicants" to "approved applicants only" and every attribution changes, because the baseline changed. Always state what your background set was.
Correlated features split credit arbitrarily. If income and account_balance are 0.9 correlated, the true joint contribution gets divided between them in a way that reflects the game-theoretic axioms rather than any causal reality. Two features that are really one signal will each look half as important as the signal is.
SHAP values are not causal. This is the most common and most damaging misreading. A SHAP value of +0.31 for income does not mean raising this person's income raises their approval probability by 0.31. It decomposes the model's output, and the model may be leaning on income as a proxy for something else entirely. Kumar et al. (2020) laid out this gap in detail. If you tell an applicant "increase your income by X and you would be approved", you are making a causal claim your explanation does not support — use a counterfactual method that actually re-queries the model instead.
SHAP tells you what the model did. It does not tell you what the world does, and it does not tell you what would happen if the applicant changed.
InterpretML: stop approximating, build a readable model
Every method above explains a black box after the fact. The alternative is a model that is accurate and directly readable, so there is no approximation error to worry about.
Explainable Boosting Machines
An EBM is a generalised additive model with pairwise interactions:
Each fi is a learned one-dimensional shape function — an arbitrary curve, not a straight line. Training uses gradient boosting in a specific way: cycle through features one at a time, fitting a tiny shallow tree on one feature per round with a very small learning rate, over thousands of rounds. Because each round touches a single feature, the effect of each feature stays separable, and the order of features stops mattering.
The payoff is that f_debt_to_income is a plottable curve giving the model's exact contribution at every value. No sampling, no kernel width, no background set. What you plot is the model.
1from interpret.glassbox import ExplainableBoostingClassifier2from interpret import show34ebm = ExplainableBoostingClassifier(interactions=10, random_state=0)5ebm.fit(X_train, y_train)67show(ebm.explain_global()) # every shape function, ranked by importance8show(ebm.explain_local(X_test[:5], y_test[:5])) # exact per-row decompositionThe pneumonia case — why readable models find real bugs
Caruana and colleagues fitted a rule-based and later an additive model to hospital data predicting death risk in pneumonia patients, to decide who could be sent home safely. The readable model contained a rule the team could see immediately: having asthma lowered predicted risk of death. It also found that a history of chest pain lowered risk.
Both patterns were real in the data and both were lethal as decision rules. Asthmatic pneumonia patients were routinely sent straight to intensive care, and the aggressive treatment gave them better outcomes. The model learned "asthma → survives" and would have recommended sending exactly those patients home. A neural network fitted to the same data almost certainly learned the same relationship — but nobody would have seen it, because there was nothing to read.
That is the argument for intrinsic interpretability in one story. The bug was not found by an accuracy metric; it was found by a clinician looking at a curve.
The rest of InterpretML
InterpretML also wraps black-box explainers behind one interface — LimeTabular, ShapKernel, PartialDependence, MorrisSensitivity — and adds performance explanations (ROC, precision–recall, confusion matrix) in the same dashboard. Having accuracy and explanation side by side matters more than it sounds: it stops people admiring a clean explanation of a model that is not actually any good.
1from interpret.blackbox import LimeTabular, ShapKernel2from interpret.perf import ROC3from interpret import show45show([LimeTabular(model.predict_proba, X_train).explain_local(X_test[:5]),6 ShapKernel(model.predict_proba, X_train[:100]).explain_local(X_test[:5]),7 ROC(model.predict_proba).explain_perf(X_test, y_test)])Choosing between them
| Method | Scope | Works on | Speed | Guarantees | Main weakness |
|---|---|---|---|---|---|
| LIME | Local | Anything with predict() | Seconds per row | None | Unstable; kernel width arbitrary |
| KernelSHAP | Local (aggregable) | Anything with predict() | Slow — minutes per row | Approximate Shapley axioms | Cost; background set sensitivity |
| TreeSHAP | Local + global | Tree ensembles only | Very fast, exact | Exact Shapley axioms | Correlated features split credit |
| EBM | Global + local | It is the model | Slower to train, instant to explain | Explanation is exact by construction | Must adopt this model class |
| Partial dependence | Global | Anything | Fast | None | Misleading under feature correlation |
| Counterfactuals | Local | Anything | Medium | Actionable by construction | May propose impossible changes |
| If your question is… | Use |
|---|---|
| "Is my model right for the wrong reason?" | Global SHAP or EBM shape functions — look for a feature that should not matter |
| "Why was this specific person rejected?" | Local SHAP waterfall, or EBM local explanation |
| "What would this person have to change?" | Counterfactuals (DiCE, Wachter et al.) — attributions cannot answer this |
| "I only have a vendor HTTP endpoint" | LIME or KernelSHAP, the only model-agnostic options |
| "A regulator will audit this and lives depend on it" | EBM or another intrinsically interpretable model |
| "Does the model treat groups differently, and why?" | Per-group SHAP summaries — compare which features drive each group's rejections |
How explainability gets misused
Explaining a model nobody validated. An explanation of a badly fitted model is a precise account of nonsense. Check performance first, per group, then explain.
Treating attribution as fairness evidence. "SHAP shows race is not in the top ten features" proves nothing. If postcode is a proxy, the influence is distributed across proxies and each one looks modest. Fairness is measured by per-group error rates, not by attribution rankings.
Fairwashing. Aïvodji et al. (2019) demonstrated that an adversary can deliberately construct a post-hoc explanation that looks reasonable while the underlying model discriminates — because the explainer is a separate, fittable object. Any explanation supplied by the party being audited should be reproducible by the auditor from the model itself.
Showing raw attributions to end users. "Your SHAP value for debt_to_income was −0.184" is not an explanation to a person who has been refused a loan. Translate to a sentence with a threshold and, where possible, an action: "your debt-to-income ratio was 0.47; applications in this band are usually declined below 0.40."
Skipping the sanity check. Before trusting any explainer on your real model, run it on a model whose behaviour you know — a linear model with fixed coefficients, or a model trained on a feature you injected as a perfect predictor. If the explainer does not recover the answer you planted, its output on the real model is not evidence of anything.
Using this on a real system
Explainability earns its keep at three specific moments, and it is worth being deliberate about which one you are in.
During development, it is a bug detector. Fit a global explanation early. Scan the top features for anything that should not be predictive: a record ID, a timestamp, a field only populated after the outcome you are predicting. Target leakage almost always shows up as one implausibly dominant feature, and it shows up in a SHAP beeswarm long before it shows up in a metric — because leakage makes the metric look better.
At review, it is the artefact that lets a domain expert object. The pneumonia rule was not caught by a data scientist; it was caught by someone who knew how asthmatic patients get triaged. Put shape functions and top features in front of that person, in their vocabulary, before launch. Budget an hour of a clinician's, underwriter's or caseworker's time — it is the highest-return hour in the project.
In production, it is a monitoring signal. Log mean absolute SHAP value per feature per week. When the ranking shifts — a feature that was fourth becomes first — something upstream changed, usually a pipeline defect or a distribution shift, and you have caught it before the accuracy metric degrades enough to trip an alarm.
One practical default: if the decision is consequential and someone will have to defend it, prefer a model you do not have to approximate. On most tabular problems an EBM is within a point or two of a gradient-boosted forest, and the difference between "here is my best estimate of what the model probably did" and "here is the model" is worth far more than two points in a room with a regulator, a journalist or a court in it.