Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

Which ML algorithms are sensitive to feature scaling?


Where units leak into the modelNeeds scalingk-NN, K-Means: distancesSVM: kernel distancesL1/L2 penalties on weightsGradient descentand neural netsPCA: variance by units
Trees are missing from this wheel because a split only asks about the order of values within one feature.

What you need to know

Why each group is sensitive

AlgorithmWhy scaling matters
k-NN, K-Means, DBSCAN, hierarchical clusteringDistance between rows is dominated by large-unit features
SVM (especially RBF kernel)The kernel and margin are computed from distances
Linear and logistic regression with L1/L2The penalty on coefficients depends on each feature's units
Any model trained with gradient descent, including neural networksDifferent ranges make the loss surface stretched and training slow or unstable
PCAIt finds directions of largest variance; a feature in big units has the most variance by definition

Why trees do not care

A decision tree asks questions like "is income greater than 7,00,000?". If you rescale income to z-scores, the same question becomes "is income greater than -0.67?". The rows on each side are exactly the same, so the tree is the same. Any transformation that keeps the order of values (scaling, shifting, even log) leaves trees unchanged.

Measuring it

The wine dataset has 13 features; one of them (proline) goes up to about 1,700, while others are below 1. Here is 5-fold cross-validated accuracy with and without standardisation.

Python
import warningsfrom sklearn.datasets import load_winefrom sklearn.ensemble import RandomForestClassifierfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import cross_val_scorefrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.pipeline import make_pipelinefrom sklearn.preprocessing import StandardScalerfrom sklearn.svm import SVCwarnings.filterwarnings("ignore")                  # unscaled logistic regression warnsX, y = load_wine(return_X_y=True)                  # 13 features; one ranges up to ~1,700models = {"k-NN": KNeighborsClassifier(), "SVM (RBF)": SVC(),          "logistic regression": LogisticRegression(max_iter=100),          "random forest": RandomForestClassifier(random_state=0)}for name, m in models.items():    raw = cross_val_score(m, X, y, cv=5).mean()    scaled = cross_val_score(make_pipeline(StandardScaler(), m), X, y, cv=5).mean()    print(f"{name:<20} raw {raw:.3f} | scaled {scaled:.3f}")
Text
k-NN                 raw 0.691 | scaled 0.949SVM (RBF)            raw 0.663 | scaled 0.983logistic regression  raw 0.956 | scaled 0.983random forest        raw 0.983 | scaled 0.983
  • k-NN and SVM jump by about 26 and 32 points. Unscaled, they are mostly measuring proline.
  • Logistic regression does fairly well unscaled but emits a "failed to converge" warning (which we silenced) and improves once scaled, because its solver reaches the optimum.
  • Random forest is exactly the same either way.

Two less obvious cases

  • Naive Bayes (Gaussian) models each feature separately, so it is not affected by scaling.
  • Plain linear regression solved exactly (ordinary least squares with no penalty) gives the same predictions with or without scaling; only the coefficients change. Scaling still helps interpretation and becomes necessary as soon as you add regularisation or use gradient descent.

A real-life example

A food-delivery team compares models for "will this order be cancelled?". An SVM performs badly and is almost dropped, until an engineer notices that order value (up to 5,000 rupees) is unscaled while "minutes since restaurant accepted" (0 to 60) and "customer's past cancellation rate" (0 to 1) are small. After adding a StandardScaler to the pipeline, the SVM matches the LightGBM model. LightGBM, being tree-based, had never needed scaling, which is why the comparison looked unfair at first.

Follow-up questions to expect

  • "Does XGBoost need scaling?" — No. Its trees split on thresholds of single features. Scaling does not hurt, but it adds a step with no benefit.
  • "Why does PCA need scaling?" — PCA picks directions with the most variance. Without scaling, a feature measured in rupees has huge variance and becomes the first component regardless of how informative it is.
  • "Do neural networks need scaling?" — Yes, in practice. Unscaled inputs make training slow and unstable, and can saturate activation functions.