Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What are common ways to reduce overfitting?


What you need to know

The toolbox, grouped by what it does

GroupTechniqueHow it helps
Check firstRemove leakage and duplicatesA leak looks like overfitting but needs a different fix
More dataCollect more rows; augment images or textNoise averages out; real patterns dominate
Less capacityFewer features, shallower trees, smaller networksLess room to memorise
ConstraintsL1/L2 regularisation, dropout, min_samples_leafPenalise or block overly specific rules
Training timeEarly stoppingStops before the model starts learning noise
AveragingBagging, random forestsMany different overfit models cancel each other's noise
Honest measurementCross-validationMakes sure an "improvement" is real

Three fixes compared

The loan-default tree from the data-splitting section memorised its training data. Here are two fixes.

Python
import numpy as npfrom sklearn.tree import DecisionTreeClassifierfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.model_selection import train_test_splitrng = np.random.default_rng(3)n = 3000income, emi_ratio, score = rng.normal(60, 20, n), rng.uniform(0, 0.8, n), rng.normal(700, 60, n)risk = -5.5 + 7 * emi_ratio - 0.03 * (score - 700) - 0.03 * (income - 60)default = rng.random(n) < 1 / (1 + np.exp(-risk))X = np.column_stack([income, emi_ratio, score])X_tr, X_te, y_tr, y_te = train_test_split(X, default, test_size=0.2, stratify=default, random_state=0)models = {    "tree, no limits": DecisionTreeClassifier(random_state=0),    "tree, min 40 rows per leaf": DecisionTreeClassifier(min_samples_leaf=40, random_state=0),    "random forest, 300 trees": RandomForestClassifier(n_estimators=300, min_samples_leaf=5, random_state=0),}for name, m in models.items():    m.fit(X_tr, y_tr)    tr, te = m.score(X_tr, y_tr), m.score(X_te, y_te)    print(f"{name:<26}: train {tr:.3f} | test {te:.3f} | gap {tr - te:.3f}")
Text
tree, no limits           : train 1.000 | test 0.832 | gap 0.168tree, min 40 rows per leaf: train 0.867 | test 0.855 | gap 0.012random forest, 300 trees  : train 0.916 | test 0.873 | gap 0.043

Requiring at least 40 applicants in every leaf stops the tree from making rules for one or two people. Training accuracy falls, test accuracy rises, and the gap almost disappears. The random forest averages 300 trees, each trained on a different sample, and gets the best test score. In both cases the training score went down and the model got better.

A sensible order

  1. Check for leakage and duplicates — cheapest, and fixes the most dramatic "overfitting".
  2. Get more or more varied data — if you can, it is the most reliable fix.
  3. Simplify — remove weak features, limit depth, shrink the network.
  4. Regularise and stop early — tune the strength on validation data.
  5. Ensemble — bagging or random forests when single models are unstable.

A real-life example

An insurance company's claim-fraud model has training AUC 0.99 and validation AUC 0.78. The team works through the list. They find 6% of claims are duplicated across splits (the same claim resubmitted) and remove them: validation AUC drops to 0.74, which is the honest number. They then drop 140 of 180 features that had almost no importance, set min_child_samples in LightGBM, and use early stopping. Validation AUC rises to 0.83, with training at 0.88. The smaller gap means the model will behave in production much as it did offline.

Follow-up questions to expect

  • "Which fix would you try first?" — Check for leakage, then more data if possible. If data is fixed, simplify and regularise, tuned on validation data.
  • "How does dropout reduce overfitting?" — During training it randomly switches off a share of neurons, so the network cannot depend on any single path and must learn more robust features.
  • "Does feature selection help?" — Yes, removing noisy or redundant features reduces the chance of matching the labels by coincidence, as long as selection is done inside cross-validation.