Course Content
Statistics & Math for AI/ML Interviews
8 sections · 30 lessons
What is correlation, and how is it quantified in ML workflows?
What you need to know
The intuition
When study hours are above average, is the exam score usually above average too? If yes, the two move together: positive correlation. If high hours go with low scores, negative correlation. If knowing one tells you nothing about the other in a straight-line way, the correlation is near zero.
The formula
Covariance multiplies each pair of deviations from the mean. Dividing by both standard deviations removes the units, so r is always between -1 and +1.
r = sum of (x - mean_x)(y - mean_y) / square root of [sum of (x - mean_x)^2 × sum of (y - mean_y)^2]Worked example by hand
Five students: study hours 1, 2, 3, 4, 5 and exam scores 52, 55, 61, 64, 68. Means are 3 and 60.
x - mean_x: -2, -1, 0, 1, 2y - mean_y: -8, -5, 1, 4, 8sum of products = 16 + 5 + 0 + 4 + 16 = 41sum of (x - mean_x)^2 = 4 + 1 + 0 + 1 + 4 = 10sum of (y - mean_y)^2 = 64 + 25 + 1 + 16 + 64 = 170r = 41 / square root of (10 × 170) = 41 / 41.23 ≈ 0.9941import pandas as pd23df = pd.DataFrame({4 "study_hours": [1, 2, 3, 4, 5],5 "exam_score": [52, 55, 61, 64, 68],6})7print(round(df["study_hours"].corr(df["exam_score"]), 3)) # Pearson8print(round(df["study_hours"].corr(df["exam_score"], method="spearman"), 3)) # Spearman0.9941.0Spearman is exactly 1 because the ranks match perfectly: more hours always means a higher score. Pearson is slightly below 1 because the points are not on a perfect straight line. On a whole DataFrame, df.corr() gives every pair at once.
Which coefficient to use
| Measure | Captures | Use when |
|---|---|---|
| Pearson | Straight-line relationships | Numeric data, roughly linear, few outliers |
| Spearman | Any monotonic relationship (always up or always down) | Curved but one-directional, or outliers |
| Kendall | Monotonic, based on pairs | Small samples, many ties |
| Cramér's V / mutual information | Any association | Categorical variables, non-linear patterns |
How it is used in ML
- Feature screening: which features move with the target?
- Redundancy: pairs with |r| above about 0.9 carry nearly the same information; keep one, or combine them.
- Leakage check: a feature with r = 0.99 to the target is suspicious — it may be computed from the target.
- Always plot. Anscombe's quartet is four small datasets with almost identical r (about 0.82) and completely different shapes, including a curve and a single outlier creating the whole correlation.
A real-life example
A food-delivery team explores its ETA data. Correlation of delivery time with distance: r = 0.78. With number of items: r = 0.12. With "restaurant rating": r = -0.05. They also find "distance_km" and "distance_road_km" have r = 0.97, so they keep only the road distance.
Then one feature, "minutes_since_order_placed_at_delivery", shows r = 0.99 with delivery time. It is the target in disguise, recorded after the delivery. Correlation analysis caught a leak before it reached training.
Follow-up questions to expect
- "What is the difference between covariance and correlation?" — Covariance has units and its size depends on scale; correlation divides by both SDs, so it is unit-free and bounded between -1 and +1.
- "Is r = 0.5 strong?" — It depends on the field; r squared = 0.25 means a straight line explains 25% of the variance in the other variable.
- "How do outliers affect Pearson r?" — A single extreme point can create or destroy a correlation; Spearman, which uses ranks, is far more robust.