Course Content
Data Science Fundamentals
4 sections · 10 lessons
Feature Scaling & Normalization
You are building a nearest-neighbour recommender. Each customer has two attributes: age in years, and annual income in pounds. To find the customer most similar to a 30-year-old earning £50,000, you compute straight-line distance.
Target: age 30, income 50,000Candidate A: age 31, income 50,000 -> distance = sqrt(1^2 + 0^2) = 1Candidate B: age 65, income 50,100 -> distance = sqrt(35^2 + 100^2) = 105.9Candidate C: age 30, income 51,000 -> distance = sqrt(0^2 + 1000^2) = 1000Candidate C is the same age as the target and earns £1,000 more. Candidate B is 35 years older. The algorithm reports that B is roughly ten times more similar to the target than C is.
That is not a subtle bias. Age has been all but erased from the calculation. The whole range of human ages is about 80 units wide; incomes span tens of thousands. Squared, a £1,000 income gap adds 1,000,000 to the squared distance, while the largest age gap possible adds about 6,400. You did not build a model that weights age lightly — you built one where age matters only when two incomes are almost identical, and nothing in the code says so.
Scaling is the fix: putting features on comparable numeric ranges so that a difference of "one unit" means something similar in each of them.
Which Algorithms Actually Care
Scaling is not universally required, and applying it blindly wastes effort and destroys interpretability. What matters is whether the algorithm compares feature magnitudes against each other.
| Algorithm | Needs scaling? | Why |
|---|---|---|
| k-nearest neighbours, k-means, DBSCAN | Essential | Built entirely on distances |
| SVM (any kernel) | Essential | Kernels are distance or dot-product based |
| PCA | Essential | Maximises variance; the largest-scale feature wins by default |
| Neural networks | Essential | Gradient descent stalls on badly conditioned inputs |
| Ridge, Lasso, ElasticNet | Essential | The penalty is applied to coefficient size, so units decide who gets penalised |
| Plain linear/logistic regression | Optional | Coefficients absorb the units; helps solver convergence and comparison |
| Decision trees, random forests, gradient boosting | No | Split on thresholds within one feature at a time; monotonic rescaling changes nothing |
| Naive Bayes | No | Models each feature's distribution separately |
The regularisation row is the one people miss. Ridge regression penalises the sum of squared coefficients. If income is in pounds its coefficient will be tiny — perhaps 0.0004 — while a coefficient on a 1-to-5 satisfaction score might be 0.8. The penalty barely touches income and crushes satisfaction, purely because of the units each was measured in. Without scaling, regularisation is applied arbitrarily.
Trees do not need scaling because they never compare one feature's magnitude with another's. Everything that measures distance, direction or coefficient size does.
Standardisation
Subtract the mean, divide by the standard deviation. The result has mean 0 and standard deviation 1, and each value expresses "how many standard deviations from average".
Worked by hand
Take five ages: 22, 28, 35, 41, 54.
mean = (22 + 28 + 35 + 41 + 54) / 5 = 36.0deviations: -14.0 -8.0 -1.0 5.0 18.0squared: 196.0 64.0 1.0 25.0 324.0sum = 610.0 variance = 610 / 4 = 152.5 sd = 12.35z-scores: 22 -> (22 - 36) / 12.35 = -1.134 28 -> (28 - 36) / 12.35 = -0.648 35 -> (35 - 36) / 12.35 = -0.081 41 -> (41 - 36) / 12.35 = 0.405 54 -> (54 - 36) / 12.35 = 1.458Note the divisor is n − 1 = 4, not 5. That is the sample standard deviation. scikit-learn's StandardScaler uses n, pandas' .std() uses n − 1, and the two therefore give slightly different z-scores on small samples. On any realistic dataset the difference is negligible, but it explains a discrepancy that otherwise looks like a bug.
1from sklearn.preprocessing import StandardScaler2import numpy as np34ages = np.array([22, 28, 35, 41, 54]).reshape(-1, 1)5scaler = StandardScaler()6print(scaler.fit_transform(ages).round(3).ravel())7print("learned mean:", scaler.mean_, "learned scale:", scaler.scale_)Standardisation does not bound the output. A value five standard deviations out becomes 5.0. It also does not change the shape of the distribution: skewed data stays exactly as skewed after standardising. This surprises people who expect scaling to "normalise" data in the sense of making it bell-shaped. It does not, and cannot — it is a shift and a stretch, nothing more.
Min-Max Normalisation
Squash everything into a fixed range, usually 0 to 1.
TextSame ages: 22, 28, 35, 41, 54. min = 22, max = 54, range = 32 22 -> (22 - 22) / 32 = 0.000 28 -> (28 - 22) / 32 = 0.188 35 -> (35 - 22) / 32 = 0.406 41 -> (41 - 22) / 32 = 0.594 54 -> (54 - 22) / 32 = 1.000Pythonfrom sklearn.preprocessing import MinMaxScalerprint(MinMaxScaler().fit_transform(ages).round(3).ravel())Where it breaks
Min-max depends on exactly two values — the smallest and the largest — so a single extreme point controls the entire transformation. Add one erroneous age of 200 to the same list:
Original Min-max without the outlier Min-max with age 200 included 22 0.000 0.000 28 0.188 0.034 35 0.406 0.073 41 0.594 0.107 54 1.000 0.180 200 — 1.000 All five real ages are now compressed into the bottom fifth of the range. The differences between them, which are the information you wanted, have been squeezed almost flat by one bad row. Standardisation on the same data would still separate the real values clearly, because the mean and standard deviation are pulled but the ordering keeps its spacing.
A second issue appears at prediction time. If a new customer is 60 years old — outside the range the scaler learned — min-max produces 1.19, breaking the guarantee that outputs lie in [0, 1]. Anything downstream that assumed the bound, such as an image pipeline or a bounded activation, now receives an out-of-range value.
Robust Scaling
Same idea as standardisation but using statistics that outliers cannot move: subtract the median, divide by the interquartile range.
x′=Q3−Q1x−medianTextAges with the erroneous 200 included: 22, 28, 35, 41, 54, 200median = (35 + 41) / 2 = 38.0Q1 = 29.75, Q3 = 50.75, IQR = 21.0 (linear interpolation, as NumPy does it) 22 -> (22 - 38) / 21 = -0.762 28 -> (28 - 38) / 21 = -0.476 35 -> (35 - 38) / 21 = -0.143 41 -> (41 - 38) / 21 = 0.143 54 -> (54 - 38) / 21 = 0.762 200 -> (200 - 38) / 21 = 7.714Pythonfrom sklearn.preprocessing import RobustScalerdata = np.array([22, 28, 35, 41, 54, 200]).reshape(-1, 1)print(RobustScaler().fit_transform(data).round(3).ravel())The five genuine ages keep their spacing, spread across roughly −0.8 to +0.8. The outlier goes to 7.7 and stays visibly an outlier rather than defining the scale for everyone else. This is the right default whenever you have not already dealt with extreme values.
Choosing Between Them
StandardScaler MinMaxScaler RobustScaler Centres on Mean Minimum Median Divides by Standard deviation Range IQR Output range Unbounded, ≈ −3 to 3 [0, 1] exactly Unbounded Outlier resistance Poor Very poor Good Preserves zero in sparse data No Yes, if min is 0 No Reach for it when Roughly symmetric data; the general default Bounded input required; known fixed range Outliers present and genuine If you want one rule: use
StandardScalerunless you have a specific reason not to, switch toRobustScalerwhen the data has real outliers, and useMinMaxScaleronly when something downstream genuinely requires a bounded input.Other transformations worth knowing
Python1from sklearn.preprocessing import PowerTransformer, QuantileTransformer, MaxAbsScaler23# Right-skewed positive data: compress the tail before scaling4df["income_log"] = np.log1p(df["income"])56# Fit an exponent that makes the data as normal as possible7yj = PowerTransformer(method="yeo-johnson") # handles zeros and negatives8X_norm = yj.fit_transform(X)910# Force any distribution onto a uniform or normal shape11qt = QuantileTransformer(output_distribution="normal", random_state=42)1213# Sparse data: scales by the largest absolute value, so zeros stay zero14X_sparse = MaxAbsScaler().fit_transform(X_sparse)The sparse-data point is practical rather than theoretical. Text data as TF-IDF vectors is typically 99.9% zeros stored in a compressed format. Subtracting a mean makes every one of those zeros non-zero, and a matrix that occupied 40 MB suddenly needs 40 GB.
MaxAbsScaleronly multiplies, so the sparsity survives.These transformations do something scaling does not: they change the shape of the distribution. A log transform turns a long right tail into something roughly symmetric. Use one when the problem is skew, and a scaler when the problem is units. They solve different things and are often applied in sequence.
The Leakage Mistake
This is the single most common serious error in applied machine learning, and it produces validation scores that look excellent and collapse in production.
Python# WRONG - scaler sees the test dataX_scaled = StandardScaler().fit_transform(X)X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2)The mean and standard deviation used to transform the training rows were computed from all the data, test rows included. Information about the test set has leaked into training. Your held-out set is no longer held out, and your estimate of performance is optimistic by an amount you cannot measure.
Python1# RIGHT - split first, learn the parameters from training data only2X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)34scaler = StandardScaler()5X_train_s = scaler.fit_transform(X_train) # fit_transform: learn, then apply6X_test_s = scaler.transform(X_test) # transform only: reuse what was learnedThe distinction between
fit_transformandtransformis the whole thing. Callingfit_transformon the test set would compute a second, different mean from the test rows — which is both leakage and inconsistency, because the two sets would then be measured on different scales.Fit on training data. Transform everything else. A test row must be processed using only knowledge available before it was seen.
The leakage becomes almost invisible during cross-validation, where the split happens repeatedly inside the loop. Pipelines solve it structurally:
Python1from sklearn.pipeline import Pipeline2from sklearn.model_selection import cross_val_score3from sklearn.neighbors import KNeighborsClassifier45pipe = Pipeline([6 ("scale", StandardScaler()),7 ("model", KNeighborsClassifier(n_neighbors=5)),8])910# The scaler is re-fitted on each training fold, never on the validation fold11scores = cross_val_score(pipe, X, y, cv=5, scoring="accuracy")12print(f"{scores.mean():.3f} +/- {scores.std():.3f}")Using a pipeline also means the scaler travels with the model. Saving
pipesaves the learned means and standard deviations alongside the estimator, so serving code cannot forget to apply them — another failure that produces confidently wrong predictions with no error message.Mixed Data Types
Real datasets contain numeric columns needing different treatment plus categorical columns needing none of it.
ColumnTransformerroutes each group to the right processing.Python1from sklearn.compose import ColumnTransformer2from sklearn.preprocessing import OneHotEncoder3from sklearn.impute import SimpleImputer4from sklearn.ensemble import GradientBoostingClassifier56symmetric = ["age", "satisfaction_score"]7skewed = ["income", "session_seconds"]8categorical = ["region", "plan"]910numeric_pipe = Pipeline([11 ("impute", SimpleImputer(strategy="median")),12 ("scale", StandardScaler()),13])1415skewed_pipe = Pipeline([16 ("impute", SimpleImputer(strategy="median")),17 ("power", PowerTransformer(method="yeo-johnson")),18])1920pre = ColumnTransformer([21 ("sym", numeric_pipe, symmetric),22 ("skew", skewed_pipe, skewed),23 ("cat", OneHotEncoder(handle_unknown="ignore"), categorical),24], remainder="drop")2526model = Pipeline([("pre", pre), ("clf", GradientBoostingClassifier())])27model.fit(X_train, y_train)Be explicit about
remainder. The default drops any column you did not list, which silently discards features if someone adds a column upstream. Setting it deliberately —"drop"or"passthrough"— makes the intent visible.What This Means When You Build Something
Return to the recommender from the opening. With
StandardScalerapplied, the three candidates change places completely. Suppose across the customer base age has mean 40 and standard deviation 14, and income has mean 45,000 and standard deviation 20,000:TextTarget: age 30 -> -0.714, income 50,000 -> 0.250Candidate A: age 31 -> -0.643, income 50,000 -> 0.250 distance = 0.071Candidate C: age 30 -> -0.714, income 51,000 -> 0.300 distance = 0.050Candidate B: age 65 -> 1.786, income 50,100 -> 0.255 distance = 2.500C is now closest, A is a whisker behind, and B — 35 years older — is far away, which is what any human would have said from the start. Same data, same algorithm, same distance metric. The only change is that both features now contribute.
Three habits keep this reliable. Scale inside a pipeline, never as a loose step, so the fit/transform boundary is enforced by structure rather than by memory. Check whether your algorithm actually needs it before reaching for a scaler — scaling inputs to a random forest wastes time and makes the model harder to interpret for no benefit. And when a distance-based or regularised model performs inexplicably badly, check the scaling first: it is a more common cause than the hyperparameters everyone tunes instead.