Course Content
Data Science Fundamentals
4 sections · 10 lessons
Correlation & Feature Importance
A team is building a churn model with 60 candidate features. To narrow things down, they compute the correlation of each feature with the target and keep anything above 0.15. One column, days_since_last_login, correlates at 0.04, so it goes in the bin.
The model underperforms. Someone eventually trains it with every feature included, and days_since_last_login comes out as the second most important predictor in the whole set.
Nothing was miscalculated. The correlation of 0.04 was correct. The problem is that Pearson correlation measures one specific thing — how well a straight line describes the relationship — and churn risk is flat for users who log in every few days, then rises sharply after about three weeks of silence. A U-shape or a hockey stick can have a correlation of essentially zero while being almost perfectly predictive. The filter did not measure importance. It measured linearity, and then that number was used as though it meant importance.
This lesson is about what correlation actually tells you, where it lies, and what to use instead when you genuinely want to know which inputs matter.
What Correlation Measures
Pearson's r is the covariance of two variables divided by the product of their standard deviations:
The division is what makes it unitless and bounded between −1 and +1. It answers exactly one question: if you drew the best straight line through this cloud of points, how tightly would the points hug it?
| Method | Measures | Data required | Robust to outliers |
|---|---|---|---|
| Pearson | Linear relationship | Numeric, roughly normal | No |
| Spearman | Monotonic relationship (rank-based) | Numeric or ordinal | Yes |
| Kendall's tau | Concordance of pairs | Numeric or ordinal, small samples | Yes |
1import pandas as pd2import numpy as np3from scipy import stats45x = df["tenure_months"]6y = df["lifetime_value"]78print("pearson :", round(x.corr(y, method="pearson"), 3))9print("spearman:", round(x.corr(y, method="spearman"), 3))Comparing the two is diagnostic. If Spearman is much higher than Pearson, the relationship is real but curved — value rises consistently with tenure, just not in a straight line. If Pearson is much higher than Spearman, a few extreme points are probably manufacturing the correlation on their own.
Reading the number honestly
| |r| | Description | Variance explained (r²) |
|---|---|---|
| 0.0 – 0.2 | Negligible | Under 4% |
| 0.2 – 0.4 | Weak | 4 – 16% |
| 0.4 – 0.6 | Moderate | 16 – 36% |
| 0.6 – 0.8 | Strong | 36 – 64% |
| 0.8 – 1.0 | Very strong | 64 – 100% |
The r² column is the honest one. A correlation of 0.5 sounds like "half", but it accounts for 25% of the variation — three quarters of what happens is driven by something else. People consistently overrate moderate correlations because 0.5 feels halfway to certain.
Significance is not strength
1r, p = stats.pearsonr(df["a"], df["b"])2n = len(df)34# Fisher z-transform confidence interval5z = np.arctanh(r)6se = 1 / np.sqrt(n - 3)7lo, hi = np.tanh([z - 1.96*se, z + 1.96*se])8print(f"r = {r:.3f}, 95% CI [{lo:.3f}, {hi:.3f}], p = {p:.3g}")With 100,000 rows, a correlation of 0.01 has a p-value of about 0.002. It is statistically significant and practically meaningless. The confidence interval is the more useful output: a correlation of 0.62 with an interval spanning 0.15 to 0.85 is a very different claim from the same 0.62 with an interval of 0.58 to 0.66.
The reason you must plot it
Anscombe's quartet is four datasets with the same mean, the same variance, and the same correlation of 0.816 — and four completely different shapes. One is a clean linear relationship. One is a perfect parabola. One is a perfect line with a single outlier dragging it. One is a vertical column of points at a single x-value plus one distant point that alone creates the entire correlation.
1import seaborn as sns2anscombe = sns.load_dataset("anscombe")3print(anscombe.groupby("dataset").apply(4 lambda g: pd.Series({"mean_x": g.x.mean(), "mean_y": g.y.mean().round(2),5 "r": g.x.corr(g.y).round(3)})6))Identical summary statistics, four incompatible stories. Only one of the four is a case where reporting r = 0.82 describes reality.
A correlation coefficient is a summary of a picture. Never report one without having looked at the picture it summarises.
Multicollinearity: Features That Duplicate Each Other
Correlation between a feature and the target is what you want. Correlation between features is what causes trouble. If height_cm and height_inches are both in your data, a linear model cannot decide how to split the credit between them. It might assign +4.2 to one and −3.9 to the other. Both coefficients are unstable, both are uninterpretable, and adding twenty more rows can flip their signs entirely.
Finding it
1corr = df[numeric_cols].corr().abs()2upper = corr.where(np.triu(np.ones(corr.shape), k=1).astype(bool))34pairs = (upper.stack()5 .sort_values(ascending=False)6 .rename("abs_r")7 .reset_index()8 .rename(columns={"level_0": "a", "level_1": "b"}))9print(pairs[pairs["abs_r"] > 0.8])Pairwise correlation misses the harder case, where no two features are strongly correlated but three together are nearly redundant — for example, when one feature is the sum of two others. The Variance Inflation Factor catches that, by regressing each feature on all the others:
Python1from statsmodels.stats.outliers_influence import variance_inflation_factor2from statsmodels.tools.tools import add_constant34X = add_constant(df[numeric_cols].dropna())5vif = pd.DataFrame({6 "feature": X.columns,7 "vif": [variance_inflation_factor(X.values, i) for i in range(X.shape[1])],8}).query("feature != 'const'").sort_values("vif", ascending=False)9print(vif)
VIF Meaning Action 1 Independent of the other features Nothing 1 – 5 Moderate overlap Usually fine 5 – 10 High Investigate > 10 Severe Drop, combine, or regularise A VIF of 10 means R² = 0.9 when that feature is predicted from the others: 90% of its information is already present elsewhere. Its coefficient's standard error is inflated by a factor of √10 ≈ 3.2, which is why the estimate jumps around.
Fixing it
Approach How Cost Drop one of the pair Keep the one that is cheaper or more interpretable Simple; may lose a little signal Combine them Average, ratio, or difference — e.g. BMI from height and weight Needs domain sense Regularise Ridge regression spreads credit across correlated features Coefficients shrink towards zero Compress Replace the correlated block with principal components New features are not interpretable An important qualifier: multicollinearity damages interpretation, not necessarily prediction. Tree-based models are largely unaffected, and a linear model's predictions can remain accurate while its individual coefficients are meaningless. If you only need forecasts, high VIF may be tolerable. If someone will read the coefficients and act on them, it is not.
Which Features Actually Matter
Linear model coefficients
Python1from sklearn.linear_model import LinearRegression2from sklearn.preprocessing import StandardScaler34Xs = StandardScaler().fit_transform(X)5model = LinearRegression().fit(Xs, y)67coefs = (pd.Series(model.coef_, index=X.columns)8 .sort_values(key=abs, ascending=False))9print(coefs)Standardising first is not optional. Without it, a coefficient of 0.003 on
income(measured in pounds) and 2,400 onsatisfaction(scored 1–5) tell you nothing about relative importance — they mostly reflect the units. After standardising, every coefficient answers the same question: how much does the target move per one-standard-deviation change in this feature?Tree-based importance
Python1from sklearn.ensemble import RandomForestRegressor23rf = RandomForestRegressor(n_estimators=300, random_state=42).fit(X, y)4imp = pd.Series(rf.feature_importances_, index=X.columns).sort_values(ascending=False)5print(imp.head(10))This is computed from how much each feature reduced impurity across all splits. It is fast and free, and it has a well-documented bias: it favours high-cardinality features. A random ID column with 50,000 distinct values will often score higher than a genuinely useful binary flag, simply because there are more places to split on it. Never take these numbers at face value without checking against a method that does not share this bias.
Permutation importance
The most trustworthy general-purpose measure. Shuffle one column, see how much worse the model gets, put it back, repeat.
Python1from sklearn.inspection import permutation_importance23r = permutation_importance(rf, X_test, y_test, n_repeats=20,4 random_state=42, scoring="r2")5perm = (pd.DataFrame({"feature": X.columns,6 "drop": r.importances_mean,7 "std": r.importances_std})8 .sort_values("drop", ascending=False))9print(perm.head(10))Two details make this work. It runs on held-out data, so it measures contribution to genuine generalisation rather than to memorisation. And it is model-agnostic — the same procedure works on a neural network, a gradient-boosted tree or a linear model, which makes results comparable across model types.
Its own weakness is correlated features. If two columns carry the same information, shuffling either one leaves the model able to recover through the other, so both score near zero and you conclude neither matters. Check the correlation matrix before believing a surprising zero.
SHAP values
Python1import shap23explainer = shap.TreeExplainer(rf)4shap_values = explainer.shap_values(X_test)56shap.summary_plot(shap_values, X_test) # per-row contributions7shap.summary_plot(shap_values, X_test, plot_type="bar") # global rankingSHAP assigns each feature a contribution to each individual prediction, with the guarantee that the contributions sum to the difference between that prediction and the average prediction. That per-row property is what the other methods lack: it lets you answer "why was this customer flagged?" rather than only "what matters on average?".
Method Speed Model-agnostic Per-prediction Main weakness Standardised coefficients Instant No — linear only No Assumes linearity Tree impurity importance Free No — trees only No Biased to high cardinality Permutation importance Moderate Yes No Confused by correlated features SHAP Slow Yes Yes Expensive; easy to over-read Multicollinearity does not necessarily hurt predictions. It destroys your ability to say what any individual coefficient means, which matters the moment someone acts on one.
Selecting Features
Three families, trading compute against quality.
Filter methods
Score each feature independently against the target, before any model exists. Fast, and they scale to thousands of columns.
Python1from sklearn.feature_selection import SelectKBest, f_regression, mutual_info_regression23# Linear association4linear_scores = SelectKBest(f_regression, k=20).fit(X, y)56# Any dependency, including non-linear - this is what the opening story needed7mi = pd.Series(mutual_info_regression(X, y, random_state=42),8 index=X.columns).sort_values(ascending=False)9print(mi.head(10))Mutual information measures how much knowing one variable reduces uncertainty about the other, in any functional form. On the churn example,
days_since_last_loginwould have scored highly on mutual information while scoring 0.04 on Pearson — the filter would have kept it.Wrapper methods
Python1from sklearn.feature_selection import RFECV2from sklearn.model_selection import KFold34selector = RFECV(RandomForestRegressor(n_estimators=100, random_state=42),5 step=1, cv=KFold(5), scoring="r2", min_features_to_select=5)6selector.fit(X, y)7print("kept:", list(X.columns[selector.support_]))Recursive feature elimination trains a model, removes the weakest feature, retrains, and repeats — with cross-validation deciding where to stop. It considers features in combination rather than one at a time, which filters cannot do, but it trains a model for every step and gets expensive quickly.
Embedded methods
Python1from sklearn.linear_model import LassoCV23lasso = LassoCV(cv=5, random_state=42).fit(Xs, y)4kept = X.columns[lasso.coef_ != 0]5print(f"{len(kept)} of {X.shape[1]} features survived")Lasso adds a penalty on the absolute size of coefficients, which drives the useless ones exactly to zero. Selection happens as a side effect of fitting, so you pay for one model rather than dozens. With correlated features, lasso tends to pick one arbitrarily and zero the rest —
ElasticNetCVhandles that case more gracefully by keeping correlated groups together.What This Means When You Build Something
Correlation and importance answer different questions, and the failure at the top of this lesson came from using one to answer the other. Correlation is a description of a pairwise relationship. Importance is a statement about a specific model's behaviour — change the model and the ranking changes, because importance was never a property of the data alone.
A workable order of operations: use the correlation matrix to find redundancy between features, not to select them; use mutual information or a quick tree model to screen for candidates, so that non-linear relationships survive the cut; use permutation importance on held-out data to decide what to keep; and use SHAP when someone needs an explanation for an individual decision.
Finally, hold onto the distinction the whole field is built on. Correlation says two things move together. It does not say one causes the other, and it never distinguishes between "A causes B", "B causes A", and "C causes both". Ice cream sales correlate with drownings because both rise in summer. If your analysis will inform a decision to change something — raise a price, send an email, alter a workflow — a correlation is a hypothesis, and only an experiment can tell you whether acting on it will do anything at all.