Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is supervised learning?


What you need to know

Inputs, labels and the mapping

Every training row has two parts:

  • Features (X) — the inputs you know at prediction time, such as a customer's tenure, bill and complaint count.
  • Label (y) — the answer you want to predict, such as "did this customer leave within 3 months?"

The algorithm searches for a function that maps X to y with as few mistakes as possible on the training rows. Because the right answers are known, the mistakes can be counted, and that count, called the loss, is what training tries to reduce.

Classification and regression

ClassificationRegression
OutputA categoryA number
ExamplesSpam or not, churn or stay, fraud or genuineFlat price, taxi fare, demand next week
Typical metricAccuracy, precision, recallMAE, RMSE

A churn model you can run

A telecom wants to know which customers will leave. It has past customers with the label "churned: yes or no".

Python
import numpy as npfrom sklearn.linear_model import LogisticRegressionfrom sklearn.model_selection import train_test_splitrng = np.random.default_rng(0)n = 2000tenure = rng.integers(1, 72, n)                 # months with the telecomcomplaints = rng.poisson(1.0, n)                # complaint calls last 90 daysbill = rng.normal(450, 150, n).round()          # monthly bill in rupeesrisk = -1.5 - 0.05 * tenure + 0.9 * complaints + 0.002 * billchurned = rng.random(n) < 1 / (1 + np.exp(-risk))   # the labelX = np.column_stack([tenure, complaints, bill])X_tr, X_te, y_tr, y_te = train_test_split(X, churned, test_size=0.25, random_state=0)model = LogisticRegression(max_iter=1000).fit(X_tr, y_tr)print("test accuracy:", round(model.score(X_te, y_te), 3))print("share who churn:", round(churned.mean(), 3))# A new customer: 3 months old, 4 complaints, 600 rupee billprint("churn probability:", model.predict_proba([[3, 4, 600]])[0, 1].round(2))
Text
test accuracy: 0.814share who churn: 0.224churn probability: 0.95

The data is simulated so the code runs anywhere. fit learns from features plus labels; predict_proba scores a customer the model has never seen. A new customer with four complaints is flagged at 0.95, so the retention team can call them first.

Notice the second line. Only 22.4% of customers churn, so a lazy model that always says "stays" would already score 77.6%. Our 81.4% beats that, but not by much. The evaluation sections of this course cover why accuracy alone can mislead.

Where supervised learning gets hard

  • Labels cost money. Someone must read 20,000 support tickets to tag them, or a doctor must read each scan.
  • Labels can be wrong or biased. If past loan officers rejected some groups unfairly, the model learns that unfairness.
  • Labels arrive late. You know whether a loan defaulted only 12 months later.

A real-life example

Production scenario. A lending app trains a default model on 80,000 past personal loans. Features are income, existing EMIs, credit score and months at current job; the label is "missed 3 or more EMIs within a year". Each new application gets a default probability in about 50 ms, and applications above a chosen risk level go to manual review.

Follow-up questions to expect

  • "What if you only have a few labels?" — Use semi-supervised learning or active learning, where the model picks the most useful rows for humans to label next. Starting from a pretrained model also reduces the labels you need.
  • "How do you know a supervised model is good?" — Evaluate it on held-out labelled data it never trained on, with a metric that matches the business cost of errors.
  • "Is supervised learning the same as classification?" — No. Classification is one kind; regression is the other.