Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

Why do we need a loss function in deep learning?


What you need to know

Three jobs the loss does

  1. Measures error — how far predictions are from targets, averaged over the batch.
  2. Provides the gradient — backpropagation computes the derivative of the loss with respect to each weight. The optimiser then moves each weight a small step in the opposite direction:
Text
w_new = w_old - learning_rate * (d loss / d w)
  1. Encodes the goal — class weights, extra penalty terms (like weight decay), or a different loss shape change what the model learns.

Why not train on accuracy directly?

Accuracy only changes when a prediction crosses the decision threshold. If the model's probability for the correct class moves from 0.30 to 0.31, accuracy does not change, so its gradient is zero almost everywhere. The optimiser gets no signal.

Cross-entropy changes smoothly: moving from 0.30 to 0.31 lowers the loss from -ln(0.30) = 1.204 to -ln(0.31) = 1.171. Every small improvement is rewarded. So we train on a smooth loss and report the metric we care about, like accuracy, recall or F1.

What a good loss needs

  • Differentiable (almost everywhere) so gradients exist. ReLU-like kinks are fine.
  • Aligned with the goal: lower loss should mean a better model for your use case.
  • Well-scaled gradients: the loss should push hard when the model is badly wrong. Cross-entropy does this for classification.

Loss vs metric

LossMetric
PurposeGuides trainingReports quality to people
Must be differentiableYesNo
ExamplesCross-entropy, MSEAccuracy, F1, recall, AUC

A real-life example

A bank trains a fraud model where 1 in 500 transactions is fraud. With plain cross-entropy, the model soon predicts "not fraud" for everything. Accuracy is 99.8%, and loss is low, because the loss treats every transaction equally and fraud is rare.

But missing a ₹50,000 fraud costs far more than reviewing a false alarm. The engineer changes the loss to match that, giving fraud examples more weight:

Python
import torch, torch.nn as nncriterion = nn.BCEWithLogitsLoss(pos_weight=torch.tensor([50.0]))logits = torch.tensor([-3.0, -3.0])   # the model says "not fraud" to both rowslabels = torch.tensor([0.0, 1.0])     # but the second row is fraudprint(criterion(logits, labels))      # the missed fraud now dominates the loss

Now a missed fraud costs 50 times more loss than a missed normal transaction. Accuracy drops slightly to 98.9%, but fraud recall rises from 0.05 to 0.82, which is what the business needed. Nothing about the architecture changed; only the loss, which is the definition of "good" the model optimises.

Follow-up questions to expect

  • "What is the difference between loss and cost function?" — Often used interchangeably. Strictly, loss is for one example and cost is the average over the dataset, sometimes plus regularisation.
  • "Can a loss go up while accuracy also goes up?" — Yes. If the model becomes more confident on wrong predictions, loss rises even as more predictions cross the threshold correctly. Watch both.
  • "What if my business metric is not differentiable?" — Train on a differentiable proxy, then tune the decision threshold or select checkpoints using the real metric on validation data.