Course Content
Machine Learning Foundations
14 sections · 70 lessons
When should you not apply feature scaling?
What you need to know
Cases where scaling is unnecessary or harmful
| Case | Why skip or change it |
|---|---|
| Tree-based models | Splits use the order of values; scaling adds a step with no benefit |
| Binary and one-hot features | Already 0/1; scaling makes them odd values like -0.4 and 2.3 that are harder to read |
| Features already on the same scale | All percentages 0–100, all pixel values 0–255: nothing to fix |
| Sparse data (text counts, TF-IDF, one-hot with many categories) | Centring destroys sparsity; use MaxAbsScaler or StandardScaler(with_mean=False) |
| Coefficients must be read in original units | Scale for training, then convert back for reporting, or skip scaling for an unregularised model |
Seeing two of these in code
1import numpy as np2from sklearn.ensemble import RandomForestClassifier3from sklearn.feature_extraction.text import TfidfVectorizer4from sklearn.preprocessing import MaxAbsScaler, StandardScaler56rng = np.random.default_rng(0)7X = np.column_stack([rng.normal(900_000, 300_000, 500), rng.uniform(21, 60, 500)])8y = (X[:, 0] < 700_000) & (X[:, 1] < 30) # a simple rule on income and age910# 1. Trees: scaling changes nothing11raw_rf = RandomForestClassifier(random_state=0).fit(X, y)12scaler = StandardScaler().fit(X)13scaled_rf = RandomForestClassifier(random_state=0).fit(scaler.transform(X), y)14same = (raw_rf.predict(X) == scaled_rf.predict(scaler.transform(X))).all()15print("random forest, identical predictions:", same)1617# 2. Sparse text features: centring would destroy sparsity18tfidf = TfidfVectorizer().fit_transform(["refund not received", "great service", "app keeps crashing"])19try:20 StandardScaler().fit(tfidf)21except ValueError as e:22 print("StandardScaler:", str(e).split(":")[0])23print("MaxAbsScaler keeps it sparse:", type(MaxAbsScaler().fit_transform(tfidf)).__name__)random forest, identical predictions: TrueStandardScaler: Cannot center sparse matricesMaxAbsScaler keeps it sparse: csr_matrixThe two forests, one on raw rupees and one on z-scores, make exactly the same predictions. For text, scikit-learn refuses to centre a sparse matrix, because a matrix that was 99.9% zeros would become 100% non-zero and might not fit in memory. MaxAbsScaler divides each column by its largest absolute value, so zeros stay zero.
The rule that always applies
When you do scale, split first, fit the scaler on the training set, and apply it to everything else. Fitting on the full dataset leaks test-set statistics into training.
A real-life example
A bank's credit team uses logistic regression and must explain to the regulator that "each extra 1 lakh of annual income lowers the odds of default by X%". A data scientist standardised all features, and the coefficients now read "per one standard deviation of income", which confuses the reviewers. The fix is not to drop scaling (the model uses L2 regularisation, which needs it) but to convert the coefficients back to original units for the report, by dividing each by the feature's standard deviation. Meanwhile, the team's LightGBM challenger model is fed raw features, because scaling would add nothing there.
Follow-up questions to expect
- "If I scale data for a random forest, is anything wrong?" — Nothing breaks and predictions stay the same; it is just an unnecessary step, and one more object to keep in sync between training and serving.
- "Should one-hot columns be scaled for a regularised linear model?" — Opinions differ. Leaving them as 0/1 is common and interpretable; scaling them makes the penalty treat rare and common categories differently. Choose one approach and apply it consistently.
- "What about scaling for gradient-boosted trees with a linear booster?" — The linear booster in XGBoost is a linear model, so scaling does matter there. The usual tree booster does not need it.