Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is a confusion matrix?
What you need to know
A classifier assigns each input to a class: fraud or genuine, spam or not spam. After you run it on a test set, every prediction falls into one of four boxes, depending on what the model said and what was actually true.
The binary layout
In scikit-learn, rows are actual and columns are predicted, with classes sorted — so for labels 0 and 1, "negative" comes first:
Predicted: 0 (No) Predicted: 1 (Yes)Actual: 0 (No) TN FPActual: 1 (Yes) FN TPSome textbooks and tools put the positive class first or swap rows and columns. Always check the axis labels before reading a matrix.
1from sklearn.metrics import confusion_matrix23actual = [0, 0, 0, 0, 0, 0, 1, 1, 1, 1] # 1 = fraud4predicted = [0, 0, 0, 0, 0, 1, 1, 1, 1, 0]56print(confusion_matrix(actual, predicted))7tn, fp, fn, tp = confusion_matrix(actual, predicted).ravel()8print(f"TN={tn} FP={fp} FN={fn} TP={tp}")[[5 1] [1 3]]TN=5 FP=1 FN=1 TP=3.ravel() flattens the 2×2 grid in the order TN, FP, FN, TP — a line worth memorising, because it is how you pull the four counts out in real code.
Every metric comes from these four numbers
accuracy = (TP + TN) / total = (3 + 5) / 10 = 0.80precision = TP / (TP + FP) = 3 / 4 = 0.75recall = TP / (TP + FN) = 3 / 4 = 0.75Once you have the matrix, you can compute any metric the interviewer asks for.
Multi-class: where the real insight is
With more than two classes, the matrix shows which pairs get confused. Here a model routes customer-support tickets into three queues:
1from sklearn.metrics import confusion_matrix23labels = ["delivery", "payment", "refund"] # support-ticket categories4actual = ["delivery"]*5 + ["payment"]*5 + ["refund"]*55predicted = (["delivery"]*4 + ["refund"] +6 ["payment"]*3 + ["refund"]*2 +7 ["payment"]*2 + ["refund"]*3)8print(confusion_matrix(actual, predicted, labels=labels))[[4 0 1] [0 3 2] [0 2 3]]"Delivery" is almost always right. But "payment" and "refund" swap with each other in both directions (the 2s off the diagonal). That is a specific, actionable finding: those two ticket types use similar words ("money", "deducted"), so the team should add features or examples that separate them. One accuracy figure (10 of 15, or 67%) could never tell you that.
A real-life example
A UPI app's fraud model is tested on 10,000 payments, of which 100 are fraud:
| Predicted genuine | Predicted fraud | |
|---|---|---|
| Actually genuine (9,900) | TN = 9,750 | FP = 150 |
| Actually fraud (100) | FN = 20 | TP = 80 |
Accuracy is 98.3%, which sounds fine but hides the detail. The matrix shows the model catches 80 of 100 frauds (recall 0.80). It also raises 230 alerts in total, and 150 of them are genuine payments wrongly blocked, so only about 1 in 3 alerts is real fraud (precision 80 / 230 ≈ 0.35). The operations head can now ask the right question: "Can my team review 230 alerts per 10,000 payments, and are 20 missed frauds acceptable?"
Follow-up questions to expect
- "How do you read a normalised confusion matrix?" — Each row is divided by its total, so the diagonal shows recall per class.
confusion_matrix(..., normalize="true")does this. - "What does the diagonal represent?" — Correct predictions. Everything off the diagonal is an error, and its position tells you which class was mistaken for which.
- "Does the confusion matrix depend on the threshold?" — Yes. It is computed from hard labels, so changing the threshold moves counts between cells.