Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is feature scaling, and why is it important?


What you need to know

The problem: units are arbitrary, numbers are not

Take two loan applicants:

Annual income (rupees)Age (years)
Applicant A9,00,00025
Applicant B9,05,00058

To a person, they are very different: one is 25, the other near retirement. Their incomes are almost the same. But to an algorithm that measures distance, the income gap is 5,000 and the age gap is 33. Income wins by a factor of about 150, and age is effectively ignored. Nothing about age became less important; its unit is just smaller.

See the damage

A k-nearest neighbours model predicts default by finding the 15 most similar past borrowers.

Python
import numpy as npfrom sklearn.model_selection import train_test_splitfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerrng = np.random.default_rng(0)n = 2000income = rng.normal(900_000, 300_000, n)      # annual income in rupeesage = rng.uniform(21, 60, n)                  # yearsemi_ratio = rng.uniform(0, 0.7, n)            # share of income on EMIs# Default depends mostly on EMI ratio and age, only a little on incomerisk = -3 + 6 * emi_ratio - 0.05 * (age - 40) - 0.000001 * (income - 900_000)default = rng.random(n) < 1 / (1 + np.exp(-risk))X = np.column_stack([income, age, emi_ratio])X_tr, X_te, y_tr, y_te = train_test_split(X, default, test_size=0.25, random_state=0)raw = KNeighborsClassifier(n_neighbors=15).fit(X_tr, y_tr)scaled = make_pipeline(StandardScaler(), KNeighborsClassifier(n_neighbors=15)).fit(X_tr, y_tr)print("k-NN on raw features   :", round(raw.score(X_te, y_te), 3))print("k-NN on scaled features:", round(scaled.score(X_te, y_te), 3))print("always predict 'repays':", round(1 - y_te.mean(), 3))
Text
k-NN on raw features   : 0.592k-NN on scaled features: 0.712always predict 'repays': 0.656

On raw features, "nearest neighbours" means "people with the most similar income", because income's numbers are huge. The EMI ratio, which ranges from 0 to 0.7 and is the real driver of default, barely affects the distance. The model does worse than always saying "repays". After standardising, all three features count, and accuracy rises to 0.712.

Why it matters, in three points

  • Distance-based models (k-NN, K-Means, SVM) compare rows by distance; big-unit features dominate.
  • Gradient descent takes much longer, or fails, when features have very different ranges (see the lesson on convergence).
  • Regularisation (L1, L2) penalises coefficient size, and coefficient size depends on units, so unscaled features are penalised unfairly.

The golden rule: fit on training data only

A scaler learns numbers from data: a mean and standard deviation, or a minimum and maximum. Those numbers are part of the model. Fit them on the training set, then reuse the same fitted scaler everywhere else. Putting the scaler inside a scikit-learn Pipeline, as above, does this automatically and saves it with the model.

A real-life example

A telecom clusters customers with K-Means on monthly data usage (in MB, up to 80,000) and number of complaint calls (0 to 6). The first result shows five clusters that differ only in data usage; complaints play no role at all, and the "unhappy heavy users" segment the team expected is invisible. After standardising both columns, a clear cluster of heavy users with frequent complaints appears: about 4% of customers, who account for a large share of churn. The retention team targets them with a network-quality callback.

Follow-up questions to expect

  • "Do all models need scaling?" — No. Tree-based models such as decision trees, random forests and gradient boosting split one feature at a time, so scaling does not change them.
  • "Why fit the scaler on training data only?" — Fitting on all data lets test-set statistics leak into training, which makes evaluation slightly optimistic and breaks the rule that test data is unseen.
  • "Should the target be scaled?" — Usually not needed for linear models or trees. For neural network regression, scaling the target can help training, but you must invert the transform on the predictions.