Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What is unsupervised learning?


What you need to know

No answer key, so what does it learn?

Without labels there is no "right answer" to compare against. The algorithm instead optimises an internal goal:

  • Clustering — put similar rows together. K-Means minimises the distance of each point to its group's centre.
  • Dimensionality reduction — keep as much information as possible in fewer columns. PCA keeps the directions with the most variance.
  • Anomaly detection — model what "normal" looks like and score how far each row is from it.
  • Association rules — find items that often appear together, such as bread and eggs.

Customer segments with K-Means

A grocery delivery app has order history but no idea which "types" of customer it has.

Python
import numpy as npfrom sklearn.cluster import KMeansfrom sklearn.preprocessing import StandardScalerrng = np.random.default_rng(1)# Shoppers: [orders per month, average order value in rupees]. No labels.X = np.vstack([    rng.normal([2, 3500], [0.8, 600], (300, 2)),   # rare, big-basket buyers    rng.normal([12, 400], [3.0, 120], (300, 2)),   # frequent, small top-ups    rng.normal([5, 1200], [1.5, 250], (300, 2)),   # regular mid-size]).clip(min=0)scaler = StandardScaler().fit(X)km = KMeans(n_clusters=3, n_init=10, random_state=0).fit(scaler.transform(X))centres = scaler.inverse_transform(km.cluster_centers_).round()for i, (orders, value) in enumerate(centres):    size = (km.labels_ == i).sum()    print(f"segment {i}: {size} shoppers, {orders:.0f} orders/month, Rs {value:.0f} per order")
Text
segment 0: 259 shoppers, 13 orders/month, Rs 394 per ordersegment 1: 297 shoppers, 2 orders/month, Rs 3517 per ordersegment 2: 344 shoppers, 5 orders/month, Rs 1113 per order

The comments in the code describe how the fake data was made, but KMeans never saw them. It found three groups on its own. A human then names them: "daily top-up", "monthly stock-up", "regular". The model gives groups; people give them meaning.

Two details matter. We scaled the features first, because order value in rupees is hundreds of times larger than order count and would otherwise decide everything. And we chose n_clusters=3 ourselves; K-Means cannot tell you the "true" number.

Evaluation without labels

  • Internal scores, such as the silhouette score, measure how tight and well separated the clusters are.
  • Stability — do similar clusters appear if you rerun on a different sample?
  • Usefulness — does the marketing team act differently for each segment, and does it pay off?

A real-life example

A bank has 2 million credit-card customers and wants to design new card offers. Nobody has labelled customers by type. The data team clusters them on spend categories, travel frequency and repayment behaviour and finds six segments, one of which is "frequent flyers who pay in full every month". The product team designs a travel card for that segment. The clustering was judged a success not by a score but by the card's sign-up rate.

Follow-up questions to expect

  • "How do you choose K in K-Means?" — Try several values and compare the elbow in within-cluster distance and the silhouette score, then check that the segments make business sense.
  • "Name some clustering algorithms besides K-Means." — DBSCAN, which finds clusters of any shape and marks noise points; hierarchical clustering; and Gaussian mixture models.
  • "Can unsupervised output feed a supervised model?" — Yes. Cluster IDs, PCA components or anomaly scores often become features for a later supervised model.