Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is classification in Machine Learning?
What you need to know
Three kinds of classification
| Kind | Meaning | Example |
|---|---|---|
| Binary | Two classes | Loan default: yes or no |
| Multi-class | One of several classes | Ticket goes to Billing, Network or Sales |
| Multi-label | Any number of labels at once | A news article tagged "cricket" and "business" |
Probabilities first, labels second
Most classifiers do not directly say "spam". They say "70% likely spam", and then a threshold decides. The default threshold is 0.5, but that is only a default.
1from sklearn.feature_extraction.text import TfidfVectorizer2from sklearn.linear_model import LogisticRegression3from sklearn.pipeline import make_pipeline45texts = [6 "Congratulations, you won a free cashback of Rs 5000, click now",7 "Your electricity connection will be cut tonight, call this number",8 "Claim your lottery prize, share bank details",9 "Free gift voucher waiting, verify your account now",10 "Team lunch moved to Friday", "Please review the attached invoice",11 "Your food order is out for delivery", "Call me when you reach the station",12]13is_spam = [1, 1, 1, 1, 0, 0, 0, 0]1415clf = make_pipeline(TfidfVectorizer(), LogisticRegression()).fit(texts, is_spam)1617msg = ["Verify your account now to claim cashback"]18p = clf.predict_proba(msg)[0]19print("classes:", clf.classes_, " probabilities:", p.round(2))20print("label at threshold 0.5:", int(p[1] >= 0.5))21print("label at threshold 0.9:", int(p[1] >= 0.9))classes: [0 1] probabilities: [0.43 0.57]label at threshold 0.5: 1label at threshold 0.9: 0The model gives 0.57 probability of spam. At the default threshold it goes to the spam folder; at 0.9 it stays in the inbox. The model did not change, only the decision rule did. A bank that fears losing a genuine customer email might pick 0.9; a phone company blocking scam SMS might pick 0.3. (With only eight training messages the probabilities are not confident; real systems train on millions.)
Algorithms and metrics
Common algorithms: logistic regression (a strong, explainable baseline), decision trees, random forests, gradient boosting, support vector machines, and neural networks for text and images.
Metrics are covered in depth later in this course. The short version: accuracy is the share of correct predictions, but if only 2% of transactions are fraud, a model that always says "genuine" is 98% accurate and useless. So you also look at precision (of those flagged, how many were right) and recall (of the real positives, how many were caught).
A real-life example
A telecom's churn model outputs a probability for each of its 5 million customers every week. The retention team can afford to call 20,000 customers. Instead of using a 0.5 threshold, they sort by probability and call the top 20,000. Classification here is really ranking by probability, and the "threshold" is set by the call-centre budget.
Follow-up questions to expect
- "How is multi-class different from multi-label?" — Multi-class picks exactly one label, and probabilities sum to 1 (softmax). Multi-label decides yes or no for each label separately (one sigmoid per label).
- "Why not always use a 0.5 threshold?" — Because the costs of the two kinds of mistake are rarely equal. Missing fraud and blocking a genuine payment cost different amounts.
- "Are the probabilities trustworthy?" — Not always. Some models, such as random forests or naive Bayes, give poorly calibrated scores; calibration methods such as
CalibratedClassifierCVfix this when the actual probability matters.