Course Content
Machine Learning Foundations
14 sections · 70 lessons
What is supervised learning?
What you need to know
Inputs, labels and the mapping
Every training row has two parts:
- Features (X) — the inputs you know at prediction time, such as a customer's tenure, bill and complaint count.
- Label (y) — the answer you want to predict, such as "did this customer leave within 3 months?"
The algorithm searches for a function that maps X to y with as few mistakes as possible on the training rows. Because the right answers are known, the mistakes can be counted, and that count, called the loss, is what training tries to reduce.
Classification and regression
| Classification | Regression | |
|---|---|---|
| Output | A category | A number |
| Examples | Spam or not, churn or stay, fraud or genuine | Flat price, taxi fare, demand next week |
| Typical metric | Accuracy, precision, recall | MAE, RMSE |
A churn model you can run
A telecom wants to know which customers will leave. It has past customers with the label "churned: yes or no".
1import numpy as np2from sklearn.linear_model import LogisticRegression3from sklearn.model_selection import train_test_split45rng = np.random.default_rng(0)6n = 20007tenure = rng.integers(1, 72, n) # months with the telecom8complaints = rng.poisson(1.0, n) # complaint calls last 90 days9bill = rng.normal(450, 150, n).round() # monthly bill in rupees10risk = -1.5 - 0.05 * tenure + 0.9 * complaints + 0.002 * bill11churned = rng.random(n) < 1 / (1 + np.exp(-risk)) # the label1213X = np.column_stack([tenure, complaints, bill])14X_tr, X_te, y_tr, y_te = train_test_split(X, churned, test_size=0.25, random_state=0)1516model = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)17print("test accuracy:", round(model.score(X_te, y_te), 3))18print("share who churn:", round(churned.mean(), 3))19# A new customer: 3 months old, 4 complaints, 600 rupee bill20print("churn probability:", model.predict_proba([[3, 4, 600]])[0, 1].round(2))test accuracy: 0.814share who churn: 0.224churn probability: 0.95The data is simulated so the code runs anywhere. fit learns from features plus labels; predict_proba scores a customer the model has never seen. A new customer with four complaints is flagged at 0.95, so the retention team can call them first.
Notice the second line. Only 22.4% of customers churn, so a lazy model that always says "stays" would already score 77.6%. Our 81.4% beats that, but not by much. The evaluation sections of this course cover why accuracy alone can mislead.
Where supervised learning gets hard
- Labels cost money. Someone must read 20,000 support tickets to tag them, or a doctor must read each scan.
- Labels can be wrong or biased. If past loan officers rejected some groups unfairly, the model learns that unfairness.
- Labels arrive late. You know whether a loan defaulted only 12 months later.
A real-life example
Production scenario. A lending app trains a default model on 80,000 past personal loans. Features are income, existing EMIs, credit score and months at current job; the label is "missed 3 or more EMIs within a year". Each new application gets a default probability in about 50 ms, and applications above a chosen risk level go to manual review.
Follow-up questions to expect
- "What if you only have a few labels?" — Use semi-supervised learning or active learning, where the model picks the most useful rows for humans to label next. Starting from a pretrained model also reduces the labels you need.
- "How do you know a supervised model is good?" — Evaluate it on held-out labelled data it never trained on, with a metric that matches the business cost of errors.
- "Is supervised learning the same as classification?" — No. Classification is one kind; regression is the other.