Course Content
Machine Learning Foundations
14 sections · 70 lessons
How do regression and classification differ in output and loss functions?
What you need to know
Output: a number versus a probability
- Regression outputs any real number: 82.6 lakh, 31 minutes, even -3 if nothing stops it.
- Classification outputs a probability between 0 and 1 for each class, then picks a class.
The final layer of a neural network reflects this:
| Task | Final layer | Output |
|---|---|---|
| Regression | Linear (no activation) | Any number |
| Binary classification | Sigmoid | One probability, 0 to 1 |
| Multi-class | Softmax | Probabilities that sum to 1 |
| Multi-label | Sigmoid per label | Independent probability for each label |
Loss: what counts as a mistake
MSE = average of (actual - predicted) squaredMAE = average of |actual - predicted|Log loss = average of -log(probability given to the true class)Here are both ideas on small numbers you can check by hand.
1from sklearn.metrics import mean_squared_error, mean_absolute_error, log_loss23# Regression: true delivery times vs predictions, in minutes4true_min = [30, 30, 30, 30]5small_misses = [33, 27, 32, 28] # every order off by 2-3 minutes6one_big_miss = [30, 30, 30, 42] # three perfect, one off by 127for name, pred in [("small misses", small_misses), ("one big miss", one_big_miss)]:8 print(f"{name}: MAE={mean_absolute_error(true_min, pred):.1f} MSE={mean_squared_error(true_min, pred):.1f}")910# Classification: the true label is "spam" (1) for one email11for p in [0.9, 0.6, 0.4, 0.01]:12 loss = log_loss([1], [p], labels=[0, 1])13 print(f"P(spam)={p:<4} -> log loss {loss:.2f}")small misses: MAE=2.5 MSE=6.5one big miss: MAE=3.0 MSE=36.0P(spam)=0.9 -> log loss 0.11P(spam)=0.6 -> log loss 0.51P(spam)=0.4 -> log loss 0.92P(spam)=0.01 -> log loss 4.61Two lessons in this output:
- MSE punishes big misses. MAE says the two sets of predictions are similar (2.5 versus 3.0). MSE says the single 12-minute miss is more than five times worse (36.0 versus 6.5), because 12 squared is 144.
- Log loss punishes confident mistakes. Saying 0.4 when the answer is spam costs 0.92. Saying 0.01, being almost certain and wrong, costs 4.61. This pushes a classifier to be honest about uncertainty.
Why not use MSE for classification?
You can compute it, but cross-entropy works better for training classifiers: it gives stronger gradients when the model is confidently wrong, and it matches the statistical meaning of predicting a probability. With a sigmoid output, MSE produces tiny gradients exactly when the model is badly wrong, so learning stalls.
Huber loss is worth naming for regression: it behaves like MSE for small errors and like MAE for large ones, so a few extreme outliers do not dominate training.
A real-life example
A credit team builds two models on the same loan applications. One predicts the loss amount if a borrower defaults (regression, trained with Huber loss because a handful of huge corporate losses would otherwise dominate MSE). The other predicts the probability of default (classification, trained with log loss). Expected loss for pricing the loan is then probability × amount. Same data, two outputs, two losses, one business number.
Follow-up questions to expect
- "When would you choose MAE over MSE as a training loss?" — When the data has outliers you do not want to dominate, or when the business cost of an error grows linearly, not with its square.
- "What does softmax do?" — It turns a list of raw scores into probabilities that are all positive and sum to 1, so the model picks one class.
- "Is loss the same as the evaluation metric?" — Not always. You might train with log loss but report recall at a fixed precision, because that is what the business cares about.