Course Content
Deep Learning Essentials
13 sections · 61 lessons
How does the choice of loss function depend on the problem type (classification vs regression)?
What you need to know
Classification losses
| Problem | Output layer | PyTorch loss |
|---|---|---|
| Binary (fraud / not fraud) | 1 logit | nn.BCEWithLogitsLoss |
| Multi-class, one label (which leaf disease) | K logits | nn.CrossEntropyLoss |
| Multi-label (several tags at once) | K logits | nn.BCEWithLogitsLoss on each |
| Heavy imbalance | same as above | add pos_weight or weight, or use focal loss |
Cross-entropy for one example is -ln(probability given to the correct class). If the model gives the right class 0.79, the loss is 0.24. If it gives 0.02, the loss is 3.9. The penalty grows very fast as the model becomes confidently wrong.
1import torch, torch.nn.functional as F23logits = torch.tensor([[2.0, 0.5, -1.0]]) # 3 classes4target = torch.tensor([0]) # correct class is 05print(logits.softmax(-1)) # tensor([[0.7856, 0.1753, 0.0391]])6print(F.cross_entropy(logits, target)) # tensor(0.2413) = -ln(0.7856)Regression losses
| Loss | Formula (per example) | Behaviour |
|---|---|---|
| MSE | (pred - y)^2 | Squares errors, so big misses dominate |
| MAE | abs(pred - y) | All errors count linearly; robust to outliers |
| Huber | squared if error is below delta, linear above | MSE near zero, MAE far out |
With errors of 2, 3 and 50 minutes in a delivery-time model, MSE averages to about 838, MAE to about 18.3, and Huber with delta=5 to about 81. The single 50-minute outlier makes up almost all of the MSE.
Why cross-entropy, not MSE, for classification
Take one example with true label 1 and a sigmoid output. When the model is confidently wrong (logit −4, probability 0.018):
MSE on the probability: gradient w.r.t. the logit = -0.035Binary cross-entropy: gradient w.r.t. the logit = -0.982MSE's gradient is tiny because the sigmoid is flat at the extremes, and MSE's derivative multiplies by that flat slope. Cross-entropy cancels the sigmoid's slope, so its gradient is simply p − y. The model learns fastest exactly when it is most wrong. Cross-entropy is also the correct loss for probabilities from a statistical view: it is the negative log-likelihood.
A real-life example
One agri-tech company, three models, three losses:
- Leaf disease classifier, one of 8 diseases per photo:
CrossEntropyLosson 8 logits. Rare diseases get higher class weights. - Leaf tagger, where one photo can show "pest damage" and "nutrient deficiency" together:
BCEWithLogitsLosson 6 independent logits. Softmax would wrongly force the tags to compete. - Yield predictor, in tonnes per hectare: first trained with MSE. A few fields had yields recorded wrongly (in kg instead of tonnes, 1,000 times too big), and those outliers dragged every prediction upward. Switching to Huber cut the average error on clean validation data by 22%, because the outliers now pushed with a linear, not squared, force.
Follow-up questions to expect
- "What is focal loss?" — Cross-entropy multiplied by
(1 − p)^gamma, which down-weights easy, already-correct examples so training focuses on hard ones. It was made for object detection, where background examples vastly outnumber objects. - "Why is MSE right for regression?" — Minimising MSE is the maximum-likelihood solution when errors are Gaussian, and it predicts the mean. MAE predicts the median.
- "What loss for ranking or embeddings?" — Contrastive or triplet losses, which pull similar pairs together and push different ones apart, or listwise ranking losses.