Course Content
Machine Learning Foundations
14 sections · 70 lessons
What kind of problems require unsupervised learning?
What you need to know
Signs a problem needs unsupervised methods
- No labels exist, and getting them would take months.
- The categories are unknown. You are asking "what groups are there?", not "which group is this?".
- The interesting cases are too rare or too new to label, such as a fraud pattern that appeared last week.
- There are too many columns to look at, and you need a compressed view.
The main problem types
| Problem | Method family | Example |
|---|---|---|
| Customer segmentation | Clustering (K-Means, DBSCAN) | Grouping shoppers by buying habits |
| Fraud or fault detection | Anomaly detection (Isolation Forest) | Odd UPI transfers, failing sensors |
| Topic discovery | Topic models, clustering of embeddings | Themes in 50,000 app reviews |
| Compression and plotting | Dimensionality reduction (PCA, UMAP) | Viewing 300 features in 2D |
| Similar items | Nearest neighbours on embeddings | "Similar products", semantic search |
Anomaly detection you can run
Fraud labels are scarce, but normal payment behaviour is plentiful. Isolation Forest learns "normal" and scores new transactions.
1import numpy as np2from sklearn.ensemble import IsolationForest34rng = np.random.default_rng(7)5# One user's UPI history: [amount in rupees, hour of day]. No fraud labels.6normal = np.column_stack([rng.gamma(2.0, 300, 500), rng.normal(14, 3, 500)])7model = IsolationForest(contamination=0.01, random_state=0).fit(normal)89new = np.array([10 [450, 13], # lunch payment, early afternoon11 [1200, 19], # groceries, evening12 [48000, 3], # large transfer at 3 a.m.13])14print(model.predict(new)) # 1 = looks normal, -1 = anomaly15print(model.score_samples(new).round(3)) # lower = more unusual[ 1 1 -1][-0.405 -0.596 -0.786]The model was never told what fraud looks like. It only learned this user's normal range and flagged the 48,000 rupee transfer at 3 a.m. as unusual. contamination=0.01 tells it to expect about 1% anomalies, which sets the cut-off. "Unusual" is not the same as "fraud": a genuine rent payment could also be flagged, so a human or a second model reviews the alerts.
A real-life example
A telecom collects 50,000 free-text complaints a month and has no categories. The team turns each complaint into an embedding (a list of numbers representing its meaning) and clusters them. The largest clusters turn out to be "network drops in a new tower area", "double-charged recharge" and "SIM activation delay". Two of these were unknown to management. The team then labels 3,000 complaints using these clusters as the category list and trains a supervised classifier to route new complaints automatically. Unsupervised learning designed the labels; supervised learning used them.
Follow-up questions to expect
- "Why not just train a fraud classifier?" — You can, when you have enough labelled fraud. Anomaly detection helps when fraud is rare, new or unlabelled, and the two are often combined.
- "How do you evaluate an anomaly detector?" — Have analysts review the top-scored cases and measure how many were real problems (precision at the top), and track known incidents it caught.
- "Is PCA a model or a preprocessing step?" — Both. It is fitted like a model, then usually used to transform data for plotting or for another model.