Data Science Fundamentals

Feature Scaling & Normalization


You are building a nearest-neighbour recommender. Each customer has two attributes: age in years, and annual income in pounds. To find the customer most similar to a 30-year-old earning £50,000, you compute straight-line distance.

Text
Target:      age 30, income 50,000Candidate A: age 31, income 50,000   -> distance = sqrt(1^2 + 0^2)       = 1Candidate B: age 65, income 50,100   -> distance = sqrt(35^2 + 100^2)    = 105.9Candidate C: age 30, income 51,000   -> distance = sqrt(0^2 + 1000^2)    = 1000

Candidate C is the same age as the target and earns £1,000 more. Candidate B is 35 years older. The algorithm reports that B is roughly ten times more similar to the target than C is.

That is not a subtle bias. Age has been all but erased from the calculation. The whole range of human ages is about 80 units wide; incomes span tens of thousands. Squared, a £1,000 income gap adds 1,000,000 to the squared distance, while the largest age gap possible adds about 6,400. You did not build a model that weights age lightly — you built one where age matters only when two incomes are almost identical, and nothing in the code says so.

Scaling is the fix: putting features on comparable numeric ranges so that a difference of "one unit" means something similar in each of them.

Which feature the distance actually listens to15 years2252.252,000 pounds4,000,0000.16DifferenceSquared, rawSquared, standardisedAgeIncomeRaw units make income nearly 18,000 times louder; standardising hands the decision back to age.
A nearest-neighbour model has no idea that pounds and years are different things — only scaling tells it.

Which Algorithms Actually Care

Scaling is not universally required, and applying it blindly wastes effort and destroys interpretability. What matters is whether the algorithm compares feature magnitudes against each other.

AlgorithmNeeds scaling?Why
k-nearest neighbours, k-means, DBSCANEssentialBuilt entirely on distances
SVM (any kernel)EssentialKernels are distance or dot-product based
PCAEssentialMaximises variance; the largest-scale feature wins by default
Neural networksEssentialGradient descent stalls on badly conditioned inputs
Ridge, Lasso, ElasticNetEssentialThe penalty is applied to coefficient size, so units decide who gets penalised
Plain linear/logistic regressionOptionalCoefficients absorb the units; helps solver convergence and comparison
Decision trees, random forests, gradient boostingNoSplit on thresholds within one feature at a time; monotonic rescaling changes nothing
Naive BayesNoModels each feature's distribution separately

The regularisation row is the one people miss. Ridge regression penalises the sum of squared coefficients. If income is in pounds its coefficient will be tiny — perhaps 0.0004 — while a coefficient on a 1-to-5 satisfaction score might be 0.8. The penalty barely touches income and crushes satisfaction, purely because of the units each was measured in. Without scaling, regularisation is applied arbitrarily.

Trees do not need scaling because they never compare one feature's magnitude with another's. Everything that measures distance, direction or coefficient size does.

Standardisation

Subtract the mean, divide by the standard deviation. The result has mean 0 and standard deviation 1, and each value expresses "how many standard deviations from average".

z=x−μσz = \frac{x - \mu}{\sigma}

Worked by hand

Take five ages: 22, 28, 35, 41, 54.

Text
mean  = (22 + 28 + 35 + 41 + 54) / 5 = 36.0deviations:  -14.0   -8.0   -1.0    5.0   18.0squared:     196.0   64.0    1.0   25.0  324.0sum = 610.0     variance = 610 / 4 = 152.5     sd = 12.35z-scores:  22 -> (22 - 36) / 12.35 = -1.134  28 -> (28 - 36) / 12.35 = -0.648  35 -> (35 - 36) / 12.35 = -0.081  41 -> (41 - 36) / 12.35 =  0.405  54 -> (54 - 36) / 12.35 =  1.458

Note the divisor is n − 1 = 4, not 5. That is the sample standard deviation. scikit-learn's StandardScaler uses n, pandas' .std() uses n − 1, and the two therefore give slightly different z-scores on small samples. On any realistic dataset the difference is negligible, but it explains a discrepancy that otherwise looks like a bug.

Python
from sklearn.preprocessing import StandardScalerimport numpy as npages = np.array([22, 28, 35, 41, 54]).reshape(-1, 1)scaler = StandardScaler()print(scaler.fit_transform(ages).round(3).ravel())print("learned mean:", scaler.mean_, "learned scale:", scaler.scale_)

Standardisation does not bound the output. A value five standard deviations out becomes 5.0. It also does not change the shape of the distribution: skewed data stays exactly as skewed after standardising. This surprises people who expect scaling to "normalise" data in the sense of making it bell-shaped. It does not, and cannot — it is a shift and a stretch, nothing more.

Min-Max Normalisation

Squash everything into a fixed range, usually 0 to 1.

x′=x−xmin⁡xmax⁡−xmin⁡x' = \frac{x - x_{\min}}{x_{\max} - x_{\min}}
Text
Same ages: 22, 28, 35, 41, 54.   min = 22, max = 54, range = 32  22 -> (22 - 22) / 32 = 0.000  28 -> (28 - 22) / 32 = 0.188  35 -> (35 - 22) / 32 = 0.406  41 -> (41 - 22) / 32 = 0.594  54 -> (54 - 22) / 32 = 1.000
Python
from sklearn.preprocessing import MinMaxScalerprint(MinMaxScaler().fit_transform(ages).round(3).ravel())

Where it breaks

Min-max depends on exactly two values — the smallest and the largest — so a single extreme point controls the entire transformation. Add one erroneous age of 200 to the same list:

OriginalMin-max without the outlierMin-max with age 200 included
220.0000.000
280.1880.034
350.4060.073
410.5940.107
541.0000.180
200—1.000

All five real ages are now compressed into the bottom fifth of the range. The differences between them, which are the information you wanted, have been squeezed almost flat by one bad row. Standardisation on the same data would still separate the real values clearly, because the mean and standard deviation are pulled but the ordering keeps its spacing.

A second issue appears at prediction time. If a new customer is 60 years old — outside the range the scaler learned — min-max produces 1.19, breaking the guarantee that outputs lie in [0, 1]. Anything downstream that assumed the bound, such as an image pipeline or a bounded activation, now receives an out-of-range value.

Robust Scaling

Same idea as standardisation but using statistics that outliers cannot move: subtract the median, divide by the interquartile range.

x′=x−medianQ3−Q1x' = \frac{x - \text{median}}{Q_3 - Q_1}

Text
Ages with the erroneous 200 included: 22, 28, 35, 41, 54, 200median = (35 + 41) / 2 = 38.0Q1 = 29.75,  Q3 = 50.75,  IQR = 21.0      (linear interpolation, as NumPy does it)  22  -> (22 - 38) / 21  = -0.762  28  -> (28 - 38) / 21  = -0.476  35  -> (35 - 38) / 21  = -0.143  41  -> (41 - 38) / 21  =  0.143  54  -> (54 - 38) / 21  =  0.762  200 -> (200 - 38) / 21 =  7.714
Python
from sklearn.preprocessing import RobustScalerdata = np.array([22, 28, 35, 41, 54, 200]).reshape(-1, 1)print(RobustScaler().fit_transform(data).round(3).ravel())

The five genuine ages keep their spacing, spread across roughly −0.8 to +0.8. The outlier goes to 7.7 and stays visibly an outlier rather than defining the scale for everyone else. This is the right default whenever you have not already dealt with extreme values.

Choosing Between Them

StandardScalerMinMaxScalerRobustScaler
Centres onMeanMinimumMedian
Divides byStandard deviationRangeIQR
Output rangeUnbounded, ≈ −3 to 3[0, 1] exactlyUnbounded
Outlier resistancePoorVery poorGood
Preserves zero in sparse dataNoYes, if min is 0No
Reach for it whenRoughly symmetric data; the general defaultBounded input required; known fixed rangeOutliers present and genuine

If you want one rule: use StandardScaler unless you have a specific reason not to, switch to RobustScaler when the data has real outliers, and use MinMaxScaler only when something downstream genuinely requires a bounded input.

Other transformations worth knowing

Python
from sklearn.preprocessing import PowerTransformer, QuantileTransformer, MaxAbsScaler# Right-skewed positive data: compress the tail before scalingdf["income_log"] = np.log1p(df["income"])# Fit an exponent that makes the data as normal as possibleyj = PowerTransformer(method="yeo-johnson")     # handles zeros and negativesX_norm = yj.fit_transform(X)# Force any distribution onto a uniform or normal shapeqt = QuantileTransformer(output_distribution="normal", random_state=42)# Sparse data: scales by the largest absolute value, so zeros stay zeroX_sparse = MaxAbsScaler().fit_transform(X_sparse)

The sparse-data point is practical rather than theoretical. Text data as TF-IDF vectors is typically 99.9% zeros stored in a compressed format. Subtracting a mean makes every one of those zeros non-zero, and a matrix that occupied 40 MB suddenly needs 40 GB. MaxAbsScaler only multiplies, so the sparsity survives.

These transformations do something scaling does not: they change the shape of the distribution. A log transform turns a long right tail into something roughly symmetric. Use one when the problem is skew, and a scaler when the problem is units. They solve different things and are often applied in sequence.

The Leakage Mistake

This is the single most common serious error in applied machine learning, and it produces validation scores that look excellent and collapse in production.

Python
# WRONG - scaler sees the test dataX_scaled = StandardScaler().fit_transform(X)X_train, X_test, y_train, y_test = train_test_split(X_scaled, y, test_size=0.2)

The mean and standard deviation used to transform the training rows were computed from all the data, test rows included. Information about the test set has leaked into training. Your held-out set is no longer held out, and your estimate of performance is optimistic by an amount you cannot measure.

Python
# RIGHT - split first, learn the parameters from training data onlyX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)scaler = StandardScaler()X_train_s = scaler.fit_transform(X_train)   # fit_transform: learn, then applyX_test_s  = scaler.transform(X_test)        # transform only: reuse what was learned

The distinction between fit_transform and transform is the whole thing. Calling fit_transform on the test set would compute a second, different mean from the test rows — which is both leakage and inconsistency, because the two sets would then be measured on different scales.

Fit on training data. Transform everything else. A test row must be processed using only knowledge available before it was seen.

The leakage becomes almost invisible during cross-validation, where the split happens repeatedly inside the loop. Pipelines solve it structurally:

Python
from sklearn.pipeline import Pipelinefrom sklearn.model_selection import cross_val_scorefrom sklearn.neighbors import KNeighborsClassifierpipe = Pipeline([    ("scale", StandardScaler()),    ("model", KNeighborsClassifier(n_neighbors=5)),])# The scaler is re-fitted on each training fold, never on the validation foldscores = cross_val_score(pipe, X, y, cv=5, scoring="accuracy")print(f"{scores.mean():.3f} +/- {scores.std():.3f}")

Using a pipeline also means the scaler travels with the model. Saving pipe saves the learned means and standard deviations alongside the estimator, so serving code cannot forget to apply them — another failure that produces confidently wrong predictions with no error message.

Mixed Data Types

Real datasets contain numeric columns needing different treatment plus categorical columns needing none of it. ColumnTransformer routes each group to the right processing.

Python
from sklearn.compose import ColumnTransformerfrom sklearn.preprocessing import OneHotEncoderfrom sklearn.impute import SimpleImputerfrom sklearn.ensemble import GradientBoostingClassifiersymmetric = ["age", "satisfaction_score"]skewed    = ["income", "session_seconds"]categorical = ["region", "plan"]numeric_pipe = Pipeline([    ("impute", SimpleImputer(strategy="median")),    ("scale", StandardScaler()),])skewed_pipe = Pipeline([    ("impute", SimpleImputer(strategy="median")),    ("power", PowerTransformer(method="yeo-johnson")),])pre = ColumnTransformer([    ("sym", numeric_pipe, symmetric),    ("skew", skewed_pipe, skewed),    ("cat", OneHotEncoder(handle_unknown="ignore"), categorical),], remainder="drop")model = Pipeline([("pre", pre), ("clf", GradientBoostingClassifier())])model.fit(X_train, y_train)

Be explicit about remainder. The default drops any column you did not list, which silently discards features if someone adds a column upstream. Setting it deliberately — "drop" or "passthrough" — makes the intent visible.

What This Means When You Build Something

Return to the recommender from the opening. With StandardScaler applied, the three candidates change places completely. Suppose across the customer base age has mean 40 and standard deviation 14, and income has mean 45,000 and standard deviation 20,000:

Text
Target:      age 30 -> -0.714,  income 50,000 -> 0.250Candidate A: age 31 -> -0.643,  income 50,000 -> 0.250   distance = 0.071Candidate C: age 30 -> -0.714,  income 51,000 -> 0.300   distance = 0.050Candidate B: age 65 ->  1.786,  income 50,100 -> 0.255   distance = 2.500

C is now closest, A is a whisker behind, and B — 35 years older — is far away, which is what any human would have said from the start. Same data, same algorithm, same distance metric. The only change is that both features now contribute.

Three habits keep this reliable. Scale inside a pipeline, never as a loose step, so the fit/transform boundary is enforced by structure rather than by memory. Check whether your algorithm actually needs it before reaching for a scaler — scaling inputs to a random forest wastes time and makes the model harder to interpret for no benefit. And when a distance-based or regularised model performs inexplicably badly, check the scaling first: it is a more common cause than the hyperparameters everyone tunes instead.