Course Content
Deep Learning Essentials
13 sections · 61 lessons
Why do we need a loss function in deep learning?
What you need to know
Three jobs the loss does
- Measures error — how far predictions are from targets, averaged over the batch.
- Provides the gradient — backpropagation computes the derivative of the loss with respect to each weight. The optimiser then moves each weight a small step in the opposite direction:
w_new = w_old - learning_rate * (d loss / d w)- Encodes the goal — class weights, extra penalty terms (like weight decay), or a different loss shape change what the model learns.
Why not train on accuracy directly?
Accuracy only changes when a prediction crosses the decision threshold. If the model's probability for the correct class moves from 0.30 to 0.31, accuracy does not change, so its gradient is zero almost everywhere. The optimiser gets no signal.
Cross-entropy changes smoothly: moving from 0.30 to 0.31 lowers the loss from -ln(0.30) = 1.204 to -ln(0.31) = 1.171. Every small improvement is rewarded. So we train on a smooth loss and report the metric we care about, like accuracy, recall or F1.
What a good loss needs
- Differentiable (almost everywhere) so gradients exist. ReLU-like kinks are fine.
- Aligned with the goal: lower loss should mean a better model for your use case.
- Well-scaled gradients: the loss should push hard when the model is badly wrong. Cross-entropy does this for classification.
Loss vs metric
| Loss | Metric | |
|---|---|---|
| Purpose | Guides training | Reports quality to people |
| Must be differentiable | Yes | No |
| Examples | Cross-entropy, MSE | Accuracy, F1, recall, AUC |
A real-life example
A bank trains a fraud model where 1 in 500 transactions is fraud. With plain cross-entropy, the model soon predicts "not fraud" for everything. Accuracy is 99.8%, and loss is low, because the loss treats every transaction equally and fraud is rare.
But missing a ₹50,000 fraud costs far more than reviewing a false alarm. The engineer changes the loss to match that, giving fraud examples more weight:
1import torch, torch.nn as nn23criterion = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([50.0]))45logits = torch.tensor([-3.0, -3.0]) # the model says "not fraud" to both rows6labels = torch.tensor([0.0, 1.0]) # but the second row is fraud7print(criterion(logits, labels)) # the missed fraud now dominates the lossNow a missed fraud costs 50 times more loss than a missed normal transaction. Accuracy drops slightly to 98.9%, but fraud recall rises from 0.05 to 0.82, which is what the business needed. Nothing about the architecture changed; only the loss, which is the definition of "good" the model optimises.
Follow-up questions to expect
- "What is the difference between loss and cost function?" — Often used interchangeably. Strictly, loss is for one example and cost is the average over the dataset, sometimes plus regularisation.
- "Can a loss go up while accuracy also goes up?" — Yes. If the model becomes more confident on wrong predictions, loss rises even as more predictions cross the threshold correctly. Watch both.
- "What if my business metric is not differentiable?" — Train on a differentiable proxy, then tune the decision threshold or select checkpoints using the real metric on validation data.