Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is the difference between normalization and standardization?
What you need to know
The formulas
Min-max normalisation: x_new = (x - min) / (max - min) -> range 0 to 1Standardisation: x_new = (x - mean) / standard_deviation -> mean 0, std 1Robust scaling: x_new = (x - median) / IQR -> median 0What an outlier does to each
Eight salaries in a small team, in thousand rupees a month. The founder earns 900.
1import numpy as np2from sklearn.preprocessing import MinMaxScaler, StandardScaler, RobustScaler34# Monthly salaries (thousand rupees) in a team of eight; the founder earns 9005salary = np.array([[35], [42], [48], [55], [60], [68], [75], [900]])67for scaler in [MinMaxScaler(), StandardScaler(), RobustScaler()]:8 out = scaler.fit_transform(salary).ravel()9 print(f"{type(scaler).__name__:<14}", np.round(out, 2))MinMaxScaler [0. 0.01 0.02 0.02 0.03 0.04 0.05 1. ]StandardScaler [-0.45 -0.42 -0.4 -0.38 -0.36 -0.33 -0.31 2.64]RobustScaler [-0.97 -0.67 -0.41 -0.11 0.11 0.45 0.75 36.24]- Min-max: the founder is 1.0 and the seven employees are squeezed between 0 and 0.05. Their real differences, 35 versus 75, are almost invisible.
- Standard: better, but still squeezed, into -0.45 to -0.31. The one outlier inflated the mean and the standard deviation.
- Robust: the seven employees spread from -0.97 to 0.75, keeping their differences clear. The founder stays an obvious outlier at 36.
The lesson: standardisation is not outlier-proof. It only avoids a hard 0–1 squeeze. For data with real outliers, use RobustScaler, cap extreme values, or apply a log transform first.
Which one to use
| Situation | Choice |
|---|---|
| Linear or logistic regression, SVM, PCA, k-NN, most tabular inputs | Standardisation (a safe default) |
| Values with natural bounds: image pixels 0–255, percentages | Min-max to 0–1 |
| Heavy outliers: income, transaction amounts, house prices | RobustScaler, or log transform then standardise |
| Sparse data such as text counts | MaxAbsScaler (keeps zeros as zeros) |
A naming warning
People use "normalisation" loosely. In scikit-learn, Normalizer is a different thing: it rescales each row to length 1, which is used for text vectors and cosine similarity. In an interview, say "min-max scaling" and "standardisation" to be precise.
A real-life example
A payments company builds a fraud model with transaction amounts from 1 rupee to 20 lakh rupees. With min-max scaling, 99% of transactions fall between 0 and 0.002, and the model struggles to tell a 200-rupee tea payment from a 20,000-rupee phone purchase. The team applies a log transform to the amount, then standardises. Now each tenfold increase in amount is an equal step, everyday payments are spread out, and the few huge transfers no longer flatten everything else.
Follow-up questions to expect
- "Does standardisation make data normally distributed?" — No. It shifts and rescales; the shape stays the same. A skewed feature stays skewed. Use a log or power transform to change the shape.
- "What does
StandardScalerstore?" — The mean and standard deviation of each feature, learned duringfiton training data. - "Which is better for neural networks?" — Both are used. Standardisation is common for tabular inputs; images are typically scaled to 0–1 or standardised per channel.