Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What does a correlation coefficient close to zero imply for feature selection?
What you need to know
Why r can be zero for a strong relationship
Pearson r only measures how well a straight line fits. Imagine customer age and churn risk: very young customers churn (they switch for offers) and very old customers churn (they find the app hard), while middle-aged customers stay. On a scatter plot this is a clear U. A straight line through a U is flat, so r ≈ 0.
1import numpy as np2from scipy.stats import pearsonr, spearmanr3from sklearn.feature_selection import mutual_info_regression45rng = np.random.default_rng(0)6age = rng.uniform(18, 70, 2_000)7churn_risk = (age - 44) ** 2 / 100 + rng.normal(0, 1, 2_000) # U-shape89print("pearson r :", round(pearsonr(age, churn_risk)[0], 3))10print("spearman :", round(spearmanr(age, churn_risk)[0], 3))11print("mutual info:", round(mutual_info_regression(age.reshape(-1, 1), churn_risk, random_state=0)[0], 3))pearson r : -0.014spearman : -0.015mutual info: 0.787Both Pearson and Spearman say "nothing here", because the relationship is not monotonic either. Mutual information, which detects any kind of dependence, is clearly positive. A decision tree would simply split young and old customers away from the middle-aged ones and use this feature heavily.
Other reasons a zero-correlation feature can matter
- Interactions. "Is weekend" might have zero correlation with delivery time on its own, but combined with "is raining" it predicts long delays. Tree models and neural networks find such combinations.
- Categorical features. City, device type or payment method encoded as numbers 1, 2, 3 give a meaningless Pearson r. Use group means, Cramér's V or MI.
- Rare but decisive values. A flag that is on for 0.1% of rows can have a tiny r and still identify most of the fraud.
Better ways to decide
- Plot the feature against the target.
- Use mutual information or Spearman as a screen.
- Use permutation importance: shuffle the feature and measure how much validation performance drops.
- Train with and without the feature and compare on a validation set.
A near-zero variance is a stronger reason to drop a feature than a near-zero correlation, though even then check rare flags.
A real-life example
An AC maker predicts service calls from outdoor temperature. The correlation across the year is close to zero, and a junior analyst suggests removing temperature. Plotting it shows a U: calls peak in the hottest weeks (units overworked) and in the coldest weeks (units switched on after months unused, with dust and gas issues). A gradient-boosted model using temperature cuts forecast error substantially compared with the model without it. The feature was essential; the statistic was the wrong tool.
Follow-up questions to expect
- "Can Spearman catch a U-shape?" — No. Spearman catches monotonic relationships only; a U goes down then up, so its rank correlation is also near zero.
- "What is permutation importance?" — Shuffle one feature's values in the validation set and measure how much the model's score drops; a big drop means the model relies on that feature.
- "Does zero correlation mean independence?" — No. Independence implies zero correlation, but zero correlation does not imply independence, as the U-shape shows.