Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do regression and classification differ in output and loss functions?


What each loss punishes hardestRegression• Output: any number, linear final layer• Loss: MSE, MAE or Huber• One 12-min miss: MSE 36 vs 6.5• Big misses cost the square of their sizeClassification• Output: probability, sigmoid or softmax• Loss: cross-entropy,also called log loss• P(spam) 0.01 on real spam: loss 4.61• Confident wrong answers cost the most
MSE punishes being far off; log loss punishes being sure and wrong — each matches what a mistake means for its output.

What you need to know

Output: a number versus a probability

  • Regression outputs any real number: 82.6 lakh, 31 minutes, even -3 if nothing stops it.
  • Classification outputs a probability between 0 and 1 for each class, then picks a class.

The final layer of a neural network reflects this:

TaskFinal layerOutput
RegressionLinear (no activation)Any number
Binary classificationSigmoidOne probability, 0 to 1
Multi-classSoftmaxProbabilities that sum to 1
Multi-labelSigmoid per labelIndependent probability for each label

Loss: what counts as a mistake

Text
MSE      = average of (actual - predicted) squaredMAE      = average of |actual - predicted|Log loss = average of -log(probability given to the true class)

Here are both ideas on small numbers you can check by hand.

Python
from sklearn.metrics import mean_squared_error, mean_absolute_error, log_loss# Regression: true delivery times vs predictions, in minutestrue_min = [30, 30, 30, 30]small_misses = [33, 27, 32, 28]      # every order off by 2-3 minutesone_big_miss = [30, 30, 30, 42]      # three perfect, one off by 12for name, pred in [("small misses", small_misses), ("one big miss", one_big_miss)]:    print(f"{name}: MAE={mean_absolute_error(true_min, pred):.1f}  MSE={mean_squared_error(true_min, pred):.1f}")# Classification: the true label is "spam" (1) for one emailfor p in [0.9, 0.6, 0.4, 0.01]:    loss = log_loss([1], [p], labels=[0, 1])    print(f"P(spam)={p:<4} -> log loss {loss:.2f}")
Text
small misses: MAE=2.5  MSE=6.5one big miss: MAE=3.0  MSE=36.0P(spam)=0.9  -> log loss 0.11P(spam)=0.6  -> log loss 0.51P(spam)=0.4  -> log loss 0.92P(spam)=0.01 -> log loss 4.61

Two lessons in this output:

  • MSE punishes big misses. MAE says the two sets of predictions are similar (2.5 versus 3.0). MSE says the single 12-minute miss is more than five times worse (36.0 versus 6.5), because 12 squared is 144.
  • Log loss punishes confident mistakes. Saying 0.4 when the answer is spam costs 0.92. Saying 0.01, being almost certain and wrong, costs 4.61. This pushes a classifier to be honest about uncertainty.

Why not use MSE for classification?

You can compute it, but cross-entropy works better for training classifiers: it gives stronger gradients when the model is confidently wrong, and it matches the statistical meaning of predicting a probability. With a sigmoid output, MSE produces tiny gradients exactly when the model is badly wrong, so learning stalls.

Huber loss is worth naming for regression: it behaves like MSE for small errors and like MAE for large ones, so a few extreme outliers do not dominate training.

A real-life example

A credit team builds two models on the same loan applications. One predicts the loss amount if a borrower defaults (regression, trained with Huber loss because a handful of huge corporate losses would otherwise dominate MSE). The other predicts the probability of default (classification, trained with log loss). Expected loss for pricing the loan is then probability × amount. Same data, two outputs, two losses, one business number.

Follow-up questions to expect

  • "When would you choose MAE over MSE as a training loss?" — When the data has outliers you do not want to dominate, or when the business cost of an error grows linearly, not with its square.
  • "What does softmax do?" — It turns a list of raw scores into probabilities that are all positive and sum to 1, so the model picks one class.
  • "Is loss the same as the evaluation metric?" — Not always. You might train with log loss but report recall at a fixed precision, because that is what the business cares about.