Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is the difference between normalization and standardization?


Min-max scaling of eight salaries0.000.010.020.020.030.040.051.000123456735k75kfounder,900kRobust scaling keeps the seven spread from -0.97 to 0.75.
One outlier squeezes seven real salary differences into five hundredths — neither min-max nor standardisation is outlier-proof.

What you need to know

The formulas

Text
Min-max normalisation:  x_new = (x - min) / (max - min)        -> range 0 to 1Standardisation:        x_new = (x - mean) / standard_deviation -> mean 0, std 1Robust scaling:         x_new = (x - median) / IQR              -> median 0

What an outlier does to each

Eight salaries in a small team, in thousand rupees a month. The founder earns 900.

Python
import numpy as npfrom sklearn.preprocessing import MinMaxScaler, StandardScaler, RobustScaler# Monthly salaries (thousand rupees) in a team of eight; the founder earns 900salary = np.array([[35], [42], [48], [55], [60], [68], [75], [900]])for scaler in [MinMaxScaler(), StandardScaler(), RobustScaler()]:    out = scaler.fit_transform(salary).ravel()    print(f"{type(scaler).__name__:<14}", np.round(out, 2))
Text
MinMaxScaler   [0.   0.01 0.02 0.02 0.03 0.04 0.05 1.  ]StandardScaler [-0.45 -0.42 -0.4  -0.38 -0.36 -0.33 -0.31  2.64]RobustScaler   [-0.97 -0.67 -0.41 -0.11  0.11  0.45  0.75 36.24]
  • Min-max: the founder is 1.0 and the seven employees are squeezed between 0 and 0.05. Their real differences, 35 versus 75, are almost invisible.
  • Standard: better, but still squeezed, into -0.45 to -0.31. The one outlier inflated the mean and the standard deviation.
  • Robust: the seven employees spread from -0.97 to 0.75, keeping their differences clear. The founder stays an obvious outlier at 36.

The lesson: standardisation is not outlier-proof. It only avoids a hard 0–1 squeeze. For data with real outliers, use RobustScaler, cap extreme values, or apply a log transform first.

Which one to use

SituationChoice
Linear or logistic regression, SVM, PCA, k-NN, most tabular inputsStandardisation (a safe default)
Values with natural bounds: image pixels 0–255, percentagesMin-max to 0–1
Heavy outliers: income, transaction amounts, house pricesRobustScaler, or log transform then standardise
Sparse data such as text countsMaxAbsScaler (keeps zeros as zeros)

A naming warning

People use "normalisation" loosely. In scikit-learn, Normalizer is a different thing: it rescales each row to length 1, which is used for text vectors and cosine similarity. In an interview, say "min-max scaling" and "standardisation" to be precise.

A real-life example

A payments company builds a fraud model with transaction amounts from 1 rupee to 20 lakh rupees. With min-max scaling, 99% of transactions fall between 0 and 0.002, and the model struggles to tell a 200-rupee tea payment from a 20,000-rupee phone purchase. The team applies a log transform to the amount, then standardises. Now each tenfold increase in amount is an equal step, everyday payments are spread out, and the few huge transfers no longer flatten everything else.

Follow-up questions to expect

  • "Does standardisation make data normally distributed?" — No. It shifts and rescales; the shape stays the same. A skewed feature stays skewed. Use a log or power transform to change the shape.
  • "What does StandardScaler store?" — The mean and standard deviation of each feature, learned during fit on training data.
  • "Which is better for neural networks?" — Both are used. Standardisation is common for tabular inputs; images are typically scaled to 0–1 or standardised per channel.