Course Content
Machine Learning Foundations
14 sections · 70 lessons
Which ML algorithms are sensitive to feature scaling?
What you need to know
Why each group is sensitive
| Algorithm | Why scaling matters |
|---|---|
| k-NN, K-Means, DBSCAN, hierarchical clustering | Distance between rows is dominated by large-unit features |
| SVM (especially RBF kernel) | The kernel and margin are computed from distances |
| Linear and logistic regression with L1/L2 | The penalty on coefficients depends on each feature's units |
| Any model trained with gradient descent, including neural networks | Different ranges make the loss surface stretched and training slow or unstable |
| PCA | It finds directions of largest variance; a feature in big units has the most variance by definition |
Why trees do not care
A decision tree asks questions like "is income greater than 7,00,000?". If you rescale income to z-scores, the same question becomes "is income greater than -0.67?". The rows on each side are exactly the same, so the tree is the same. Any transformation that keeps the order of values (scaling, shifting, even log) leaves trees unchanged.
Measuring it
The wine dataset has 13 features; one of them (proline) goes up to about 1,700, while others are below 1. Here is 5-fold cross-validated accuracy with and without standardisation.
1import warnings2from sklearn.datasets import load_wine3from sklearn.ensemble import RandomForestClassifier4from sklearn.linear_model import LogisticRegression5from sklearn.model_selection import cross_val_score6from sklearn.neighbors import KNeighborsClassifier7from sklearn.pipeline import make_pipeline8from sklearn.preprocessing import StandardScaler9from sklearn.svm import SVC1011warnings.filterwarnings("ignore") # unscaled logistic regression warns12X, y = load_wine(return_X_y=True) # 13 features; one ranges up to ~1,70013models = {"k-NN": KNeighborsClassifier(), "SVM (RBF)": SVC(),14 "logistic regression": LogisticRegression(max_iter=100),15 "random forest": RandomForestClassifier(random_state=0)}16for name, m in models.items():17 raw = cross_val_score(m, X, y, cv=5).mean()18 scaled = cross_val_score(make_pipeline(StandardScaler(), m), X, y, cv=5).mean()19 print(f"{name:<20} raw {raw:.3f} | scaled {scaled:.3f}")k-NN raw 0.691 | scaled 0.949SVM (RBF) raw 0.663 | scaled 0.983logistic regression raw 0.956 | scaled 0.983random forest raw 0.983 | scaled 0.983- k-NN and SVM jump by about 26 and 32 points. Unscaled, they are mostly measuring proline.
- Logistic regression does fairly well unscaled but emits a "failed to converge" warning (which we silenced) and improves once scaled, because its solver reaches the optimum.
- Random forest is exactly the same either way.
Two less obvious cases
- Naive Bayes (Gaussian) models each feature separately, so it is not affected by scaling.
- Plain linear regression solved exactly (ordinary least squares with no penalty) gives the same predictions with or without scaling; only the coefficients change. Scaling still helps interpretation and becomes necessary as soon as you add regularisation or use gradient descent.
A real-life example
A food-delivery team compares models for "will this order be cancelled?". An SVM performs badly and is almost dropped, until an engineer notices that order value (up to 5,000 rupees) is unscaled while "minutes since restaurant accepted" (0 to 60) and "customer's past cancellation rate" (0 to 1) are small. After adding a StandardScaler to the pipeline, the SVM matches the LightGBM model. LightGBM, being tree-based, had never needed scaling, which is why the comparison looked unfair at first.
Follow-up questions to expect
- "Does XGBoost need scaling?" — No. Its trees split on thresholds of single features. Scaling does not hurt, but it adds a step with no benefit.
- "Why does PCA need scaling?" — PCA picks directions with the most variance. Without scaling, a feature measured in rupees has huge variance and becomes the first component regardless of how informative it is.
- "Do neural networks need scaling?" — Yes, in practice. Unscaled inputs make training slow and unstable, and can saturate activation functions.