Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do you decide whether the issue is data, features, or model choice?


Validation accuracy as training rows grow0.840.860.850.880.850.890.850.91logistic regressionrandom forest320 rows960 rows1,920 rows3,200 rowsLogistic is flat: more rows will not help. The forest is still climbing: more data will.
A flat curve says the limit is the model or the features; a rising one says the limit is data.

What you need to know

Three possible bottlenecks, three different fixes:

Data volume

  • Validation still rising as rows are added
  • Big train–validation gap
  • Fix: collect more labelled rows, augment

Features (signal)

  • Validation flat at a poor score
  • Even a flexible model does not help
  • Fix: new features, better domain inputs

Model choice

  • Simple model flat, flexible model much higher
  • Train and validation equal, both mediocre
  • Fix: switch to a model that can express the pattern

The learning curve

A learning curve plots training and validation scores against the number of training rows. scikit-learn computes it for you:

Python
import numpy as npfrom sklearn.datasets import make_classificationfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import learning_curveX, y = make_classification(n_samples=4000, n_features=20, n_informative=8,                           n_clusters_per_class=3, flip_y=0.05, random_state=0)sizes = [0.1, 0.3, 0.6, 1.0]for name, model in [("logistic", LogisticRegression()),                    ("forest", RandomForestClassifier(random_state=0, n_jobs=-1))]:    n, train, val = learning_curve(model, X, y, train_sizes=sizes, cv=5)    print(name)    for rows, t, v in zip(n, train.mean(axis=1), val.mean(axis=1)):        print(f"  {rows:5d} rows  train={t:.2f}  validation={v:.2f}")
Text
logistic    320 rows  train=0.86  validation=0.84    960 rows  train=0.85  validation=0.85   1920 rows  train=0.85  validation=0.85   3200 rows  train=0.85  validation=0.85forest    320 rows  train=1.00  validation=0.86    960 rows  train=1.00  validation=0.88   1920 rows  train=1.00  validation=0.89   3200 rows  train=1.00  validation=0.91

Read the two models separately.

  • Logistic regression is flat at 0.85 from 960 rows onward, and training and validation are equal. It has learned everything a straight-line model can from these features. More data will not help. The limit is the model's shape — confirmed because the forest, on the same features, does better.
  • Random forest is still climbing (0.86 → 0.91) and has a large gap to its training score. It is data-limited: more rows would likely push it higher.

If both models had been flat at 0.85, the bottleneck would be the features: the information simply is not there, and the fix is new inputs, not new algorithms.

Label noise sets a ceiling

If human labellers agree with each other only 90% of the time — common for tasks like "is this review abusive?" — then no model can reliably exceed about 90% agreement with those labels. Measure agreement on a sample of double-labelled rows before blaming the model.

A quick checklist

  1. Train a strong, flexible model (gradient boosting) with defaults. If it is much better than your model → model choice.
  2. If it is not better → features or labels.
  3. Plot a learning curve for the best model. Still rising → data volume. Flat → features.
  4. Check label agreement → sets the realistic ceiling.

A real-life example

A hospital builds a model to predict 30-day readmission from discharge records. Logistic regression gets ROC-AUC 0.68; gradient boosting gets 0.69. Since the flexible model barely helps, the problem is not model choice. The learning curve with 40,000 patients is flat after 15,000, so more patients of the same kind would not help either. The team concludes the features are the limit: discharge records miss what happens at home. They add pharmacy refill data and whether a follow-up visit was booked, and ROC-AUC rises to 0.76.

Follow-up questions to expect

  • "How do you get more data if you can't label more?" — Augmentation for images and text, weak labels from rules, semi-supervised learning, or active learning to label only the most informative rows.
  • "How would you check whether a feature has any signal?" — Compare validation scores with and without it, or use permutation importance on validation data.
  • "What if training score is low even on a flexible model?" — Suspect a bug: labels shuffled relative to features, a broken join, or all-null columns.