Machine Learning Essentials

Supervised vs Unsupervised Learning


Two spreadsheets sit on your desk. They came from the same bank, on the same day, exported by the same analyst.

The first has 5,000 rows, one per loan the bank issued three years ago. Columns: applicant age, income, employment length, existing debt, credit score — and then one final column, repaid, containing yes or no. Every row's outcome is known, because three years have passed and the loans have either been paid off or written off.

The second has 5,000 rows too, one per current customer. Columns: age, account balance, number of products held, monthly transaction count, branch visits per year. That is all. There is no final column. Nothing in the file says what any of these customers are.

These two files demand completely different machinery. Not different algorithms in the sense that a hammer and a mallet are different — different in the sense that one lets you ask "what will happen?" and the other only lets you ask "what is going on here?". The gap between them is the oldest and most useful division in machine learning, and almost every practical mistake beginners make is a mistake about which side of it they are standing on.

The label column is the whole distinctionSupervised: an answer key exists• One column holds the known outcome• Regression predicts a number• Classification picks one of a fixed set• Accuracy can be measured against truthUnsupervised: no answer key• Only inputs, no outcome column• Clustering groups rows that resemble• Reduction mergescolumns saying the same thing• Results can be judged, never marked
Both spreadsheets came from the same bank; only one of them carries the column that says what happened.

The label is the entire distinction

A label (also called the target, the outcome, or y) is a column containing the answer you eventually want to predict for new, unseen rows. In the first spreadsheet, repaid is a label. In the second spreadsheet there is no label at all.

Everything follows from that:

Supervised learning learns from examples where the correct answer is attached. Unsupervised learning looks for structure in data where no correct answer exists.

The word "supervised" is a metaphor about a teacher marking homework. You hand the algorithm a stack of worked problems with the answers written at the bottom. It makes a guess on each one, compares its guess to the real answer, sees exactly how wrong it was, and adjusts. That feedback loop — guess, compare, adjust — is only possible because someone already knew the answers.

Remove the answer key and that loop is gone. An unsupervised algorithm can never be told "you got that one wrong", because there is nothing to be wrong about. It can only report structure it found: these rows resemble each other, this one is unlike all the rest, these five columns are mostly repeating the same information.

Why people get this backwards

The common error is assuming the label is whatever column looks most interesting. It is not. The label is defined by the question you will ask at prediction time.

Suppose you want to predict which customers will churn next quarter. You have a table of customers with a churned column. That looks supervised — and it is, but only if churned was recorded before you would have needed the prediction. If the column was filled in by looking at the same quarter you are trying to predict, you have built a model that requires the future to predict the future. It will score beautifully in testing and be useless in production. This failure is called leakage, and it is the single most common reason a model with 99% accuracy on your laptop performs at chance level in the real world.

Supervised learning, and the two shapes it comes in

Supervised problems split by what kind of thing the label is.

Regression: the label is a number on a continuous scale

Predicting a house price, a delivery time in minutes, tomorrow's temperature, a patient's blood pressure. The defining property is that the answers have an order and the gaps between them mean something: predicting £310,000 when the truth was £300,000 is a better outcome than predicting £500,000, and "better" here is measurable.

Because errors are measurable in the units of the thing itself, regression models are scored by how far off they are. A model that predicts house prices with an average error of £8,400 is straightforwardly better than one averaging £21,000.

Classification: the label is one of a fixed set of categories

Spam or not spam. Repaid or defaulted. Which of seven product categories this support ticket belongs to. Here the categories usually have no meaningful ordering — "spam" is not larger than "not spam" — so the errors are counted, not measured. You got 43 wrong out of 500, and the interesting question becomes which kind of wrong.

RegressionClassification
Label looks like3.7, 412000, −11.2"spam", "churn", class 3
Number of possible answersInfiniteFixed and finite
Typical questionHow much? How many?Which one? Yes or no?
Scored bySize of the error (MAE, RMSE)Count of errors, by type (precision, recall)
Being "slightly wrong"Meaningful and commonNot a thing — you are right or wrong
Example algorithmsLinear regression, gradient boosting regressorsLogistic regression, random forests, k-NN

The trap between them

Some labels look numeric but are categories, and some look categorical but are numeric. Star ratings from 1 to 5 are the classic ambiguity. Treated as regression, a model predicting 4.3 stars is being sensible — it thinks the review is a bit better than a 4. Treated as classification with five classes, the model has no idea that class 5 is closer to class 4 than to class 1, so confusing a 5 with a 1 costs exactly the same as confusing a 5 with a 4.

Neither framing is wrong. But choosing regression when the business genuinely needs a hard category, or classification when the ordering carries real information, quietly costs you accuracy that no amount of tuning recovers.

What supervision costs

Supervised learning is more powerful and more directly useful. It is also more expensive, and the expense is almost entirely in the labels.

  • Someone has to produce them. A radiologist marking 10,000 scans is not a data pipeline problem, it is a payroll problem.
  • They may not exist yet. To label loan defaults you must wait for loans to default. A three-year loan gives you a three-year lag between collecting data and being able to train on it.
  • They are noisy. Human labellers disagree. On sentiment tasks, two trained annotators typically agree only about 80–85% of the time. Your model cannot be more accurate than your labels are consistent, so an 85%-consistent label set puts a hard ceiling near 85% on any honest evaluation.

Your model's accuracy ceiling is set by your label quality, not your algorithm. Spending a week improving labels routinely beats spending a week tuning hyperparameters.

Unsupervised learning: structure without an answer key

Return to the second spreadsheet — 5,000 customers, no outcome column. You cannot predict anything, because nothing has been labelled as worth predicting. What you can do is ask what shape the data has.

Clustering: which rows belong together

A clustering algorithm groups rows so that rows within a group resemble each other more than they resemble rows in other groups. Run it on the customer file and you might get four groups: high-balance customers who barely transact, low-balance customers who transact constantly, customers who hold many products, and a small group who visit branches often and do little else.

Note carefully what happened. The algorithm did not name those groups. It produced numbers — cluster 0, cluster 1, cluster 2, cluster 3 — and a human looked at the averages inside each and decided what to call them. Interpretation is always a human step. The algorithm found that four bundles exist; it has no concept of "wealthy" or "dormant".

Dimensionality reduction: which columns are secretly the same column

If your dataset has monthly_income, annual_income, and weekly_income, three columns are carrying one column's worth of information. Real datasets do this constantly in subtler ways: height and weight, ad clicks and ad impressions, dozens of survey questions that all measure roughly one underlying attitude.

Dimensionality reduction compresses many correlated columns into a few new ones that retain most of the variation. A 200-column dataset might compress to 20 columns holding 95% of the original variation — which makes models faster to train, easier to visualise, and often more accurate, because the discarded 5% was largely noise.

Anomaly detection: which rows do not belong anywhere

Learn what normal looks like from unlabelled data, then flag whatever sits far from it. This is how card fraud systems work in practice, because the alternative — supervised classification — needs labelled examples of fraud, and new fraud patterns are by definition ones you have never labelled.

Association rules: what co-occurs

Given millions of shopping baskets, find the combinations that appear together far more often than chance would predict. "Customers who buy nappies also buy baby wipes" is trivial. "Customers who buy premium dog food also buy houseplants" is the kind of finding that changes a shelf layout.

The hard part: unsupervised results cannot be marked

This deserves its own section because it catches everyone.

With a supervised model, evaluation is unambiguous. Hold back 1,000 rows the model never saw, predict them, compare to the known answers, report the error. If your model gets 940 of 1,000 right, that number means something to anyone.

With clustering, there is no held-back truth. Suppose you cluster customers into 4 groups and someone else clusters them into 7. Which is right? There is no answer in the data. Both are valid descriptions. You can compute internal measures — how tight the clusters are, how well separated — and those help, but they measure geometric neatness, not usefulness.

SupervisedUnsupervised
InputFeatures + labelsFeatures only
GoalPredict the label for new rowsDescribe the structure of existing rows
Feedback during trainingExplicit error signal per exampleNone
EvaluationObjective, on held-out dataIndirect; needs human judgement
"Correct answer" exists?YesNo
Main costObtaining labelsValidating that the output is useful
Fails quietly whenLabels leak the futureFeatures are on wildly different scales

An unsupervised result is a hypothesis, not a conclusion. It tells you "these rows group together"; only a human, or a downstream experiment, can tell you whether that grouping is worth acting on.

The scale trap in that last row is worth spelling out, because it silently ruins more clustering projects than any other single mistake. Most clustering measures distance between rows. If income ranges from 20,000 to 200,000 and number_of_products ranges from 1 to 5, then a £10,000 income difference contributes 10,000 units of distance while the entire product range contributes 4. The algorithm has effectively clustered on income alone and ignored everything else — and it will not warn you. Standardising every feature to comparable ranges before clustering is not optional.

The middle ground: semi-supervised and self-supervised

The real world often hands you 500 labelled rows and 50,000 unlabelled ones, because labelling is expensive but collecting raw data is nearly free.

Semi-supervised learning uses both. One simple approach: train on the 500 labelled rows, predict the 50,000, keep only predictions the model is highly confident about, add those to the training set as if they were real labels, and retrain. This works when it works and quietly amplifies your model's own biases when it does not — a confidently wrong prediction becomes a confidently wrong training example.

Self-supervised learning is a cleverer trick: manufacture labels from the data's own structure. Take a sentence, hide one word, and train a model to predict the hidden word. No human labelled anything, yet the training loop is fully supervised because the answer was always there — you just covered it up. Every large language model is trained this way, which is why the internet is enough training data.

Choosing, in practice

Work through this in order.

  1. Write down the sentence you want the system to output. "This applicant will default" is a prediction. "Our customers fall into these groups" is a description. Predictions are supervised; descriptions are unsupervised.
  2. If it is a prediction, check the label actually exists — recorded, for enough rows, at a point in time strictly before the moment of prediction. If it does not exist, your project's first phase is a labelling project, not a modelling one.
  3. If the label is a number on a scale, that is regression. If it is a category, that is classification.
  4. If there is no label and no way to get one, ask what description would be useful: groups (clustering), fewer columns (dimensionality reduction), or odd rows (anomaly detection).

The same data, both ways

Here is the split in code. Notice that the supervised call takes two arguments and the unsupervised call takes one — that single difference in the API is the distinction.

Python
import pandas as pdfrom sklearn.model_selection import train_test_splitfrom sklearn.ensemble import RandomForestClassifierfrom sklearn.cluster import KMeansfrom sklearn.preprocessing import StandardScalerfrom sklearn.metrics import accuracy_scoredf = pd.read_csv("customers.csv")features = ["age", "balance", "num_products", "transactions_per_month"]# ---- Supervised: we have a 'churned' column ----X = df[features]y = df["churned"]                      # the answer keyX_train, X_test, y_train, y_test = train_test_split(    X, y, test_size=0.2, random_state=42, stratify=y)clf = RandomForestClassifier(n_estimators=200, random_state=42)clf.fit(X_train, y_train)              # TWO arguments: questions and answerspreds = clf.predict(X_test)print("accuracy:", accuracy_score(y_test, preds))   # objective score# ---- Unsupervised: pretend 'churned' does not exist ----X_scaled = StandardScaler().fit_transform(df[features])   # scaling is mandatory herekm = KMeans(n_clusters=4, n_init=10, random_state=42)km.fit(X_scaled)                       # ONE argument: no answers to givedf["segment"] = km.labels_print(df.groupby("segment")[features].mean())   # a human reads this and names the groups

The supervised branch ends with a number you can put in a report. The unsupervised branch ends with a table someone has to interpret. That asymmetry never goes away.

Using both together, which is what real systems do

The framing so far has been "pick one". In production they are usually stacked.

A fraud team runs anomaly detection over all transactions to surface the few thousand strangest ones per day. Analysts review a sample of those and label them fraud or not. Those labels train a supervised classifier. The classifier catches known fraud patterns cheaply and at scale; the anomaly detector keeps catching the new patterns the classifier has never seen. Neither alone is sufficient.

Similarly, dimensionality reduction is most often used not as an end in itself but as a preprocessing step: compress 300 noisy columns to 30 informative ones, then train a supervised model on those. The unsupervised step never predicts anything — it just makes the supervised step work better.

What this means when you build something

Before writing a line of modelling code, answer one question in writing: at prediction time, on a row I have never seen, what exactly am I asking this system to output?

If you can complete that sentence with a specific value, and you can point to a column in your historical data containing that value, recorded before the moment of prediction, you have a supervised problem and you should immediately go and check the quality and consistency of that column — because it caps everything.

If you cannot complete the sentence, you do not yet have a machine learning problem. You have an exploration problem, and unsupervised methods are the right tools, with the understanding that their output is a starting point for a human conversation rather than a deliverable in itself.

The projects that fail are rarely the ones that pick the wrong algorithm. They are the ones that spend three weeks building a supervised model on a label nobody validated, or that deliver a cluster analysis nobody can act on because no one agreed in advance what a useful grouping would look like. Get the framing right and the algorithm choice is a detail. Get it wrong and no algorithm saves you.