Course Content
Machine Learning Foundations
14 sections · 70 lessons
What are common ways to reduce overfitting?
What you need to know
The toolbox, grouped by what it does
| Group | Technique | How it helps |
|---|---|---|
| Check first | Remove leakage and duplicates | A leak looks like overfitting but needs a different fix |
| More data | Collect more rows; augment images or text | Noise averages out; real patterns dominate |
| Less capacity | Fewer features, shallower trees, smaller networks | Less room to memorise |
| Constraints | L1/L2 regularisation, dropout, min_samples_leaf | Penalise or block overly specific rules |
| Training time | Early stopping | Stops before the model starts learning noise |
| Averaging | Bagging, random forests | Many different overfit models cancel each other's noise |
| Honest measurement | Cross-validation | Makes sure an "improvement" is real |
Three fixes compared
The loan-default tree from the data-splitting section memorised its training data. Here are two fixes.
1import numpy as np2from sklearn.tree import DecisionTreeClassifier3from sklearn.ensemble import RandomForestClassifier4from sklearn.model_selection import train_test_split56rng = np.random.default_rng(3)7n = 30008income, emi_ratio, score = rng.normal(60, 20, n), rng.uniform(0, 0.8, n), rng.normal(700, 60, n)9risk = -5.5 + 7 * emi_ratio - 0.03 * (score - 700) - 0.03 * (income - 60)10default = rng.random(n) < 1 / (1 + np.exp(-risk))11X = np.column_stack([income, emi_ratio, score])12X_tr, X_te, y_tr, y_te = train_test_split(X, default, test_size=0.2, stratify=default, random_state=0)1314models = {15 "tree, no limits": DecisionTreeClassifier(random_state=0),16 "tree, min 40 rows per leaf": DecisionTreeClassifier(min_samples_leaf=40, random_state=0),17 "random forest, 300 trees": RandomForestClassifier(n_estimators=300, min_samples_leaf=5, random_state=0),18}19for name, m in models.items():20 m.fit(X_tr, y_tr)21 tr, te = m.score(X_tr, y_tr), m.score(X_te, y_te)22 print(f"{name:<26}: train {tr:.3f} | test {te:.3f} | gap {tr - te:.3f}")tree, no limits : train 1.000 | test 0.832 | gap 0.168tree, min 40 rows per leaf: train 0.867 | test 0.855 | gap 0.012random forest, 300 trees : train 0.916 | test 0.873 | gap 0.043Requiring at least 40 applicants in every leaf stops the tree from making rules for one or two people. Training accuracy falls, test accuracy rises, and the gap almost disappears. The random forest averages 300 trees, each trained on a different sample, and gets the best test score. In both cases the training score went down and the model got better.
A sensible order
- Check for leakage and duplicates — cheapest, and fixes the most dramatic "overfitting".
- Get more or more varied data — if you can, it is the most reliable fix.
- Simplify — remove weak features, limit depth, shrink the network.
- Regularise and stop early — tune the strength on validation data.
- Ensemble — bagging or random forests when single models are unstable.
A real-life example
An insurance company's claim-fraud model has training AUC 0.99 and validation AUC 0.78. The team works through the list. They find 6% of claims are duplicated across splits (the same claim resubmitted) and remove them: validation AUC drops to 0.74, which is the honest number. They then drop 140 of 180 features that had almost no importance, set min_child_samples in LightGBM, and use early stopping. Validation AUC rises to 0.83, with training at 0.88. The smaller gap means the model will behave in production much as it did offline.
Follow-up questions to expect
- "Which fix would you try first?" — Check for leakage, then more data if possible. If data is fixed, simplify and regularise, tuned on validation data.
- "How does dropout reduce overfitting?" — During training it randomly switches off a share of neurons, so the network cannot depend on any single path and must learn more robust features.
- "Does feature selection help?" — Yes, removing noisy or redundant features reduces the chance of matching the labels by coincidence, as long as selection is done inside cross-validation.