Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is a confusion matrix?


10,000 UPI payments, 100 of them fraudTN 9,750FP 150FN 20TP 80Predicted genuinePredicted fraudActual genuineActual fraudAccuracy 98.3 percent, recall 0.80, precision 80 of 230 alerts.
The same four cells give every metric, but only the matrix shows that two in three alerts are false alarms.

What you need to know

A classifier assigns each input to a class: fraud or genuine, spam or not spam. After you run it on a test set, every prediction falls into one of four boxes, depending on what the model said and what was actually true.

The binary layout

In scikit-learn, rows are actual and columns are predicted, with classes sorted — so for labels 0 and 1, "negative" comes first:

Text
                 Predicted: 0 (No)   Predicted: 1 (Yes)Actual: 0 (No)        TN                  FPActual: 1 (Yes)       FN                  TP

Some textbooks and tools put the positive class first or swap rows and columns. Always check the axis labels before reading a matrix.

Python
from sklearn.metrics import confusion_matrixactual    = [0, 0, 0, 0, 0, 0, 1, 1, 1, 1]   # 1 = fraudpredicted = [0, 0, 0, 0, 0, 1, 1, 1, 1, 0]print(confusion_matrix(actual, predicted))tn, fp, fn, tp = confusion_matrix(actual, predicted).ravel()print(f"TN={tn} FP={fp} FN={fn} TP={tp}")
Text
[[5 1] [1 3]]TN=5 FP=1 FN=1 TP=3

.ravel() flattens the 2×2 grid in the order TN, FP, FN, TP — a line worth memorising, because it is how you pull the four counts out in real code.

Every metric comes from these four numbers

Text
accuracy  = (TP + TN) / total                 = (3 + 5) / 10 = 0.80precision = TP / (TP + FP)                    = 3 / 4       = 0.75recall    = TP / (TP + FN)                    = 3 / 4       = 0.75

Once you have the matrix, you can compute any metric the interviewer asks for.

Multi-class: where the real insight is

With more than two classes, the matrix shows which pairs get confused. Here a model routes customer-support tickets into three queues:

Python
from sklearn.metrics import confusion_matrixlabels = ["delivery", "payment", "refund"]           # support-ticket categoriesactual    = ["delivery"]*5 + ["payment"]*5 + ["refund"]*5predicted = (["delivery"]*4 + ["refund"] +             ["payment"]*3 + ["refund"]*2 +             ["payment"]*2 + ["refund"]*3)print(confusion_matrix(actual, predicted, labels=labels))
Text
[[4 0 1] [0 3 2] [0 2 3]]

"Delivery" is almost always right. But "payment" and "refund" swap with each other in both directions (the 2s off the diagonal). That is a specific, actionable finding: those two ticket types use similar words ("money", "deducted"), so the team should add features or examples that separate them. One accuracy figure (10 of 15, or 67%) could never tell you that.

A real-life example

A UPI app's fraud model is tested on 10,000 payments, of which 100 are fraud:

Predicted genuinePredicted fraud
Actually genuine (9,900)TN = 9,750FP = 150
Actually fraud (100)FN = 20TP = 80

Accuracy is 98.3%, which sounds fine but hides the detail. The matrix shows the model catches 80 of 100 frauds (recall 0.80). It also raises 230 alerts in total, and 150 of them are genuine payments wrongly blocked, so only about 1 in 3 alerts is real fraud (precision 80 / 230 ≈ 0.35). The operations head can now ask the right question: "Can my team review 230 alerts per 10,000 payments, and are 20 missed frauds acceptable?"

Follow-up questions to expect

  • "How do you read a normalised confusion matrix?" — Each row is divided by its total, so the diagonal shows recall per class. confusion_matrix(..., normalize="true") does this.
  • "What does the diagonal represent?" — Correct predictions. Everything off the diagonal is an error, and its position tells you which class was mistaken for which.
  • "Does the confusion matrix depend on the threshold?" — Yes. It is computed from hard labels, so changing the threshold moves counts between cells.