Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is classification in Machine Learning?


Five messages, one model, two thresholds0.120.350.570.810.9501234spam at 0.5,inbox at 0.9spam at bothScores are P(spam); the window is what a 0.5 threshold sends to the spam folder.
The model outputs the same probabilities either way — the threshold is a business decision layered on top.

What you need to know

Three kinds of classification

KindMeaningExample
BinaryTwo classesLoan default: yes or no
Multi-classOne of several classesTicket goes to Billing, Network or Sales
Multi-labelAny number of labels at onceA news article tagged "cricket" and "business"

Probabilities first, labels second

Most classifiers do not directly say "spam". They say "70% likely spam", and then a threshold decides. The default threshold is 0.5, but that is only a default.

Python
from sklearn.feature_extraction.text import TfidfVectorizerfrom sklearn.linear_model import LogisticRegressionfrom sklearn.pipeline import make_pipelinetexts = [    "Congratulations, you won a free cashback of Rs 5000, click now",    "Your electricity connection will be cut tonight, call this number",    "Claim your lottery prize, share bank details",    "Free gift voucher waiting, verify your account now",    "Team lunch moved to Friday", "Please review the attached invoice",    "Your food order is out for delivery", "Call me when you reach the station",]is_spam = [1, 1, 1, 1, 0, 0, 0, 0]clf = make_pipeline(TfidfVectorizer(), LogisticRegression()).fit(texts, is_spam)msg = ["Verify your account now to claim cashback"]p = clf.predict_proba(msg)[0]print("classes:", clf.classes_, " probabilities:", p.round(2))print("label at threshold 0.5:", int(p[1] >= 0.5))print("label at threshold 0.9:", int(p[1] >= 0.9))
Text
classes: [0 1]  probabilities: [0.43 0.57]label at threshold 0.5: 1label at threshold 0.9: 0

The model gives 0.57 probability of spam. At the default threshold it goes to the spam folder; at 0.9 it stays in the inbox. The model did not change, only the decision rule did. A bank that fears losing a genuine customer email might pick 0.9; a phone company blocking scam SMS might pick 0.3. (With only eight training messages the probabilities are not confident; real systems train on millions.)

Algorithms and metrics

Common algorithms: logistic regression (a strong, explainable baseline), decision trees, random forests, gradient boosting, support vector machines, and neural networks for text and images.

Metrics are covered in depth later in this course. The short version: accuracy is the share of correct predictions, but if only 2% of transactions are fraud, a model that always says "genuine" is 98% accurate and useless. So you also look at precision (of those flagged, how many were right) and recall (of the real positives, how many were caught).

A real-life example

A telecom's churn model outputs a probability for each of its 5 million customers every week. The retention team can afford to call 20,000 customers. Instead of using a 0.5 threshold, they sort by probability and call the top 20,000. Classification here is really ranking by probability, and the "threshold" is set by the call-centre budget.

Follow-up questions to expect

  • "How is multi-class different from multi-label?" — Multi-class picks exactly one label, and probabilities sum to 1 (softmax). Multi-label decides yes or no for each label separately (one sigmoid per label).
  • "Why not always use a 0.5 threshold?" — Because the costs of the two kinds of mistake are rarely equal. Missing fraud and blocking a genuine payment cost different amounts.
  • "Are the probabilities trustworthy?" — Not always. Some models, such as random forests or naive Bayes, give poorly calibrated scores; calibration methods such as CalibratedClassifierCV fix this when the actual probability matters.