Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are ensemble learning methods?
What you need to know
Why combining models works
If several models make errors that are not perfectly correlated, averaging them cancels some errors. Suppose three models each have 80% accuracy and make mistakes independently. A majority vote is right when at least two are right:
P(all 3 right) = 0.8^3 = 0.512P(exactly 2) = 3 * 0.8^2 * 0.2 = 0.384Majority right = 0.896In practice models' errors are correlated, so the gain is smaller, but the principle holds: diversity between models is what makes an ensemble useful.
The main methods
| Method | How | Reduces | Examples |
|---|---|---|---|
| Bagging | Train models in parallel on random resamples of data, then average | Variance | Random forest |
| Boosting | Train models in sequence, each focusing on previous errors | Bias (and variance) | AdaBoost, gradient boosting, XGBoost, LightGBM, CatBoost |
| Stacking | Train a meta-model on the base models' predictions | Both | Kaggle-style blends |
| Voting / averaging | Majority vote or mean of probabilities | Variance | Any set of models |
Ensembles in deep learning
- Seed ensembles — train the same network 3 to 5 times with different random seeds and data order, then average their probabilities.
- Snapshot ensembles — save checkpoints at several points in one training run (often with a cyclic learning rate) and average them.
- Test-time augmentation (TTA) — predict on several flipped or cropped versions of an input and average.
- Weight averaging — average the weights of checkpoints (SWA, model soups) to get ensemble-like gains at single-model cost.
- Dropout is sometimes described as an implicit ensemble of thinned networks.
Averaging probabilities from several PyTorch models is one line:
probs = torch.stack([m(x).softmax(dim=-1) for m in models]).mean(dim=0)pred = probs.argmax(dim=-1)The costs
- Inference cost and latency multiply by the number of models.
- More models to maintain, version and monitor.
- Harder to explain a decision.
A common production pattern is distillation: train one small "student" model to copy the ensemble's outputs, then deploy only the student.
A real-life example
A crop-disease competition team has three models on their validation set: EfficientNet at 91.2%, ResNet-50 at 90.4%, and a vision transformer at 90.8%. Their errors overlap only partly: the CNNs confuse two blights with similar colour, while the transformer confuses leaf mould with dust on the leaf.
Averaging the three models' probabilities gives 93.1%. Adding horizontal-flip TTA gives 93.6%. For the competition, that is worth it.
For the farmer app, three models would triple inference time on a phone. So the team distils the ensemble into one EfficientNet student, which reaches 92.5%: most of the gain at a third of the cost.
Follow-up questions to expect
- "Bagging vs boosting in one line?" — Bagging trains independent models in parallel to reduce variance; boosting trains dependent models in sequence to reduce bias.
- "Why does random forest also sample features?" — To make the trees more different from each other. Less correlated trees mean more errors cancel.
- "When is an ensemble not worth it?" — When latency or cost limits are tight, when the models are nearly identical so their errors do not cancel, or when explainability is required.