Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

How does the choice of loss function depend on the problem type (classification vs regression)?


Gradient on the logit when the true label is 10.018−0.035−0.9820.269−0.288−0.7310.881−0.025−0.119pMSE gradBCE gradlogit −4logit −1logit 2
Cross-entropy pushes hardest when the model is most confidently wrong, while MSE through a sigmoid barely pushes at all.

What you need to know

Classification losses

ProblemOutput layerPyTorch loss
Binary (fraud / not fraud)1 logitnn.BCEWithLogitsLoss
Multi-class, one label (which leaf disease)K logitsnn.CrossEntropyLoss
Multi-label (several tags at once)K logitsnn.BCEWithLogitsLoss on each
Heavy imbalancesame as aboveadd pos_weight or weight, or use focal loss

Cross-entropy for one example is -ln(probability given to the correct class). If the model gives the right class 0.79, the loss is 0.24. If it gives 0.02, the loss is 3.9. The penalty grows very fast as the model becomes confidently wrong.

Python
import torch, torch.nn.functional as Flogits = torch.tensor([[2.0, 0.5, -1.0]])   # 3 classestarget = torch.tensor([0])                  # correct class is 0print(logits.softmax(-1))                   # tensor([[0.7856, 0.1753, 0.0391]])print(F.cross_entropy(logits, target))      # tensor(0.2413) = -ln(0.7856)

Regression losses

LossFormula (per example)Behaviour
MSE(pred - y)^2Squares errors, so big misses dominate
MAEabs(pred - y)All errors count linearly; robust to outliers
Hubersquared if error is below delta, linear aboveMSE near zero, MAE far out

With errors of 2, 3 and 50 minutes in a delivery-time model, MSE averages to about 838, MAE to about 18.3, and Huber with delta=5 to about 81. The single 50-minute outlier makes up almost all of the MSE.

Why cross-entropy, not MSE, for classification

Take one example with true label 1 and a sigmoid output. When the model is confidently wrong (logit −4, probability 0.018):

Text
MSE on the probability:  gradient w.r.t. the logit = -0.035Binary cross-entropy:    gradient w.r.t. the logit = -0.982

MSE's gradient is tiny because the sigmoid is flat at the extremes, and MSE's derivative multiplies by that flat slope. Cross-entropy cancels the sigmoid's slope, so its gradient is simply p − y. The model learns fastest exactly when it is most wrong. Cross-entropy is also the correct loss for probabilities from a statistical view: it is the negative log-likelihood.

A real-life example

One agri-tech company, three models, three losses:

  • Leaf disease classifier, one of 8 diseases per photo: CrossEntropyLoss on 8 logits. Rare diseases get higher class weights.
  • Leaf tagger, where one photo can show "pest damage" and "nutrient deficiency" together: BCEWithLogitsLoss on 6 independent logits. Softmax would wrongly force the tags to compete.
  • Yield predictor, in tonnes per hectare: first trained with MSE. A few fields had yields recorded wrongly (in kg instead of tonnes, 1,000 times too big), and those outliers dragged every prediction upward. Switching to Huber cut the average error on clean validation data by 22%, because the outliers now pushed with a linear, not squared, force.

Follow-up questions to expect

  • "What is focal loss?" — Cross-entropy multiplied by (1 − p)^gamma, which down-weights easy, already-correct examples so training focuses on hard ones. It was made for object detection, where background examples vastly outnumber objects.
  • "Why is MSE right for regression?" — Minimising MSE is the maximum-likelihood solution when errors are Gaussian, and it predicts the mean. MAE predicts the median.
  • "What loss for ranking or embeddings?" — Contrastive or triplet losses, which pull similar pairs together and push different ones apart, or listwise ranking losses.