Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is underfitting?
What you need to know
The signature
| Training score | Validation score | Gap | |
|---|---|---|---|
| Underfitting | Low | Low | Small |
| Good fit | High | High | Small |
| Overfitting | Very high | Clearly lower | Large |
The key point: in underfitting, even training performance is poor. The model cannot fit the data it already saw.
Watch it happen
Food orders per minute across a day have two peaks, lunch and dinner. A straight line cannot bend twice.
1import numpy as np2from sklearn.linear_model import LinearRegression3from sklearn.ensemble import GradientBoostingRegressor4from sklearn.model_selection import train_test_split56rng = np.random.default_rng(1)7hour = rng.uniform(0, 24, 3000)8# Orders per minute in a city: a lunch peak at 13:00 and a bigger dinner peak at 20:309orders = (40 * np.exp(-((hour - 13) / 1.5) ** 2) + 60 * np.exp(-((hour - 20.5) / 1.8) ** 2)10 + 5 + rng.normal(0, 4, 3000))11X = hour.reshape(-1, 1)12X_tr, X_te, y_tr, y_te = train_test_split(X, orders, test_size=0.25, random_state=0)1314for name, model in [("straight line ", LinearRegression()),15 ("gradient boosting", GradientBoostingRegressor(random_state=0))]:16 model.fit(X_tr, y_tr)17 print(f"{name}: train R2 {model.score(X_tr, y_tr):.3f} | test R2 {model.score(X_te, y_te):.3f}")straight line : train R2 0.335 | test R2 0.315gradient boosting: train R2 0.958 | test R2 0.946The straight line explains only about a third of the variation, on training data too. The gap is tiny (0.335 versus 0.315), which tells you it is not overfitting; the model is simply too weak. Gradient boosting can model two peaks and scores about 0.95 on both. Adding more rows would not help the line; its shape is the limit.
Common causes and fixes
| Cause | Fix |
|---|---|
| Model too simple (straight line on curved data) | Use a more flexible model, such as trees or boosting |
| Missing information (no "hour of day" feature) | Add better features |
| Too much regularisation | Reduce the penalty strength |
| Training stopped too early | Train longer, or raise the learning rate carefully |
| Tree depth or leaf size too restrictive | Relax max_depth, min_samples_leaf |
Sometimes the problem is not the model at all. If the features simply do not contain the answer, such as predicting churn from customer ID only, no model will do well. That is a data problem, not an underfitting problem.
A real-life example
A telecom predicts monthly data usage with linear regression on age and plan price. Training and validation R² are both around 0.30. The team first tries a bigger model; it barely helps, reaching 0.34. Then they add features that actually drive usage: hours of video streamed last month, device type, whether the customer is on a 5G handset. R² rises to 0.78 on both sets. The underfitting was mostly missing information, not a weak algorithm.
Follow-up questions to expect
- "How do you tell underfitting from overfitting?" — Look at the training score. Low training and low validation means underfitting; high training with much lower validation means overfitting.
- "Will more data fix underfitting?" — Usually not. The model cannot use the data it has. Add capacity or better features instead.
- "Can a deep neural network underfit?" — Yes, if it is trained too briefly, with a bad learning rate, or with too much regularisation or dropout.