Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

What kind of problems require unsupervised learning?


Unsupervised learning designs the labels50,000untaggedcomplaintsEmbed eachcomplaintas numbersCluster: networkdrops, double chargeLabel 3,000using thoseclustersTrain asupervisedrouter
Clustering answered 'what categories exist?', which is the question supervised learning cannot ask.

What you need to know

Signs a problem needs unsupervised methods

  • No labels exist, and getting them would take months.
  • The categories are unknown. You are asking "what groups are there?", not "which group is this?".
  • The interesting cases are too rare or too new to label, such as a fraud pattern that appeared last week.
  • There are too many columns to look at, and you need a compressed view.

The main problem types

ProblemMethod familyExample
Customer segmentationClustering (K-Means, DBSCAN)Grouping shoppers by buying habits
Fraud or fault detectionAnomaly detection (Isolation Forest)Odd UPI transfers, failing sensors
Topic discoveryTopic models, clustering of embeddingsThemes in 50,000 app reviews
Compression and plottingDimensionality reduction (PCA, UMAP)Viewing 300 features in 2D
Similar itemsNearest neighbours on embeddings"Similar products", semantic search

Anomaly detection you can run

Fraud labels are scarce, but normal payment behaviour is plentiful. Isolation Forest learns "normal" and scores new transactions.

Python
import numpy as npfrom sklearn.ensemble import IsolationForestrng = np.random.default_rng(7)# One user's UPI history: [amount in rupees, hour of day]. No fraud labels.normal = np.column_stack([rng.gamma(2.0, 300, 500), rng.normal(14, 3, 500)])model = IsolationForest(contamination=0.01, random_state=0).fit(normal)new = np.array([    [450, 13],     # lunch payment, early afternoon    [1200, 19],    # groceries, evening    [48000, 3],    # large transfer at 3 a.m.])print(model.predict(new))              # 1 = looks normal, -1 = anomalyprint(model.score_samples(new).round(3))  # lower = more unusual
Text
[ 1  1 -1][-0.405 -0.596 -0.786]

The model was never told what fraud looks like. It only learned this user's normal range and flagged the 48,000 rupee transfer at 3 a.m. as unusual. contamination=0.01 tells it to expect about 1% anomalies, which sets the cut-off. "Unusual" is not the same as "fraud": a genuine rent payment could also be flagged, so a human or a second model reviews the alerts.

A real-life example

A telecom collects 50,000 free-text complaints a month and has no categories. The team turns each complaint into an embedding (a list of numbers representing its meaning) and clusters them. The largest clusters turn out to be "network drops in a new tower area", "double-charged recharge" and "SIM activation delay". Two of these were unknown to management. The team then labels 3,000 complaints using these clusters as the category list and trains a supervised classifier to route new complaints automatically. Unsupervised learning designed the labels; supervised learning used them.

Follow-up questions to expect

  • "Why not just train a fraud classifier?" — You can, when you have enough labelled fraud. Anomaly detection helps when fraud is rare, new or unlabelled, and the two are often combined.
  • "How do you evaluate an anomaly detector?" — Have analysts review the top-scored cases and measure how many were real problems (precision at the top), and track known incidents it caught.
  • "Is PCA a model or a preprocessing step?" — Both. It is fitted like a model, then usually used to transform data for plotting or for another model.