Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do you decide whether the issue is data, features, or model choice?
What you need to know
Three possible bottlenecks, three different fixes:
Data volume
- Validation still rising as rows are added
- Big train–validation gap
- Fix: collect more labelled rows, augment
Features (signal)
- Validation flat at a poor score
- Even a flexible model does not help
- Fix: new features, better domain inputs
Model choice
- Simple model flat, flexible model much higher
- Train and validation equal, both mediocre
- Fix: switch to a model that can express the pattern
The learning curve
A learning curve plots training and validation scores against the number of training rows. scikit-learn computes it for you:
1import numpy as np2from sklearn.datasets import make_classification3from sklearn.ensemble import RandomForestClassifier4from sklearn.linear_model import LogisticRegression5from sklearn.model_selection import learning_curve67X, y = make_classification(n_samples=4000, n_features=20, n_informative=8,8 n_clusters_per_class=3, flip_y=0.05, random_state=0)9sizes = [0.1, 0.3, 0.6, 1.0]10for name, model in [("logistic", LogisticRegression()),11 ("forest", RandomForestClassifier(random_state=0, n_jobs=-1))]:12 n, train, val = learning_curve(model, X, y, train_sizes=sizes, cv=5)13 print(name)14 for rows, t, v in zip(n, train.mean(axis=1), val.mean(axis=1)):15 print(f" {rows:5d} rows train={t:.2f} validation={v:.2f}")logistic 320 rows train=0.86 validation=0.84 960 rows train=0.85 validation=0.85 1920 rows train=0.85 validation=0.85 3200 rows train=0.85 validation=0.85forest 320 rows train=1.00 validation=0.86 960 rows train=1.00 validation=0.88 1920 rows train=1.00 validation=0.89 3200 rows train=1.00 validation=0.91Read the two models separately.
- Logistic regression is flat at 0.85 from 960 rows onward, and training and validation are equal. It has learned everything a straight-line model can from these features. More data will not help. The limit is the model's shape — confirmed because the forest, on the same features, does better.
- Random forest is still climbing (0.86 → 0.91) and has a large gap to its training score. It is data-limited: more rows would likely push it higher.
If both models had been flat at 0.85, the bottleneck would be the features: the information simply is not there, and the fix is new inputs, not new algorithms.
Label noise sets a ceiling
If human labellers agree with each other only 90% of the time — common for tasks like "is this review abusive?" — then no model can reliably exceed about 90% agreement with those labels. Measure agreement on a sample of double-labelled rows before blaming the model.
A quick checklist
- Train a strong, flexible model (gradient boosting) with defaults. If it is much better than your model → model choice.
- If it is not better → features or labels.
- Plot a learning curve for the best model. Still rising → data volume. Flat → features.
- Check label agreement → sets the realistic ceiling.
A real-life example
A hospital builds a model to predict 30-day readmission from discharge records. Logistic regression gets ROC-AUC 0.68; gradient boosting gets 0.69. Since the flexible model barely helps, the problem is not model choice. The learning curve with 40,000 patients is flat after 15,000, so more patients of the same kind would not help either. The team concludes the features are the limit: discharge records miss what happens at home. They add pharmacy refill data and whether a follow-up visit was booked, and ROC-AUC rises to 0.76.
Follow-up questions to expect
- "How do you get more data if you can't label more?" — Augmentation for images and text, weak labels from rules, semi-supervised learning, or active learning to label only the most informative rows.
- "How would you check whether a feature has any signal?" — Compare validation scores with and without it, or use permutation importance on validation data.
- "What if training score is low even on a flexible model?" — Suspect a bug: labels shuffled relative to features, a broken join, or all-null columns.