Machine Learning Foundations

Course Content

Machine Learning Foundations

14 sections · 70 lessons

How do labels change the learning process in ML?


The same half-moon data, two signalsWith labels• Loss compares eachprediction to the truth• k-NN learns the curved boundary• Test accuracy 1.0• Success means correctWithout labels• Goal is distance to a cluster centre• K-Means splits into two round blobs• Agrees with trueclasses 78.5% of the time• Success means useful, judged by people
K-Means did exactly what it was asked; without labels it simply had no way to know the classes were curved.

What you need to know

Labels define what "good" means

In supervised learning, training is a loop: predict, compare with the label, measure the error, adjust. The loss function turns "compare with the label" into a number.

Without labels, there is nothing to compare against. K-Means minimises the distance from each point to its cluster centre; PCA maximises the variance kept. These goals describe the data's shape. They do not know what you actually care about.

The same data, with and without labels

Here two classes form interlocking half-moons. A supervised model uses the labels; K-Means only sees the points.

Python
from sklearn.datasets import make_moonsfrom sklearn.cluster import KMeansfrom sklearn.neighbors import KNeighborsClassifierfrom sklearn.metrics import accuracy_score# Two classes shaped like interlocking half-moonsX, y = make_moons(n_samples=600, noise=0.08, random_state=0)X_tr, y_tr, X_te, y_te = X[:400], y[:400], X[400:], y[400:]# With labels: learn the boundary the labels defineknn = KNeighborsClassifier(n_neighbors=5).fit(X_tr, y_tr)print("supervised accuracy:", accuracy_score(y_te, knn.predict(X_te)))# Without labels: KMeans groups points by distance to a centregroups = KMeans(n_clusters=2, n_init=10, random_state=0).fit_predict(X_te)match = max(accuracy_score(y_te, groups), accuracy_score(y_te, 1 - groups))print("clusters agree with true classes:", round(match, 3))
Text
supervised accuracy: 1.0clusters agree with true classes: 0.785

With labels, the model learns the curved boundary exactly. Without them, K-Means splits the points into two round blobs because that is what its goal rewards, and it agrees with the real classes only 78.5% of the time. (We check both label orders because cluster numbers 0 and 1 are arbitrary.) Same data, different signal, different result.

Label quality is a ceiling

  • Noisy labels. If 10% of "spam" tags are users misclicking, the model learns some wrong lessons, and your test labels are noisy too, so even the evaluation is blurred.
  • Biased labels. If past loan approvals were unfair to one group, "approved" labels encode that unfairness.
  • Ambiguous definitions. If one annotator calls a message "complaint" and another calls it "query", the model cannot beat their disagreement.

A real-life example

Production scenario. An e-commerce firm trained a product-review sentiment model on star ratings: 4–5 stars "positive", 1–2 "negative". Accuracy plateaued at 84%. An audit of 500 reviews found many 1-star reviews were about late delivery, not the product, and many 5-star reviews contained complaints. The labels measured delivery mood, not product opinion. Relabelling 8,000 reviews with a clear guide raised accuracy more than any change to the model had.

Follow-up questions to expect

  • "What can you do about noisy labels?" — Write a clear labelling guide, have two people label a sample and measure their agreement, relabel the rows where the model and label disagree most, and use a clean test set.
  • "What is weak supervision?" — Creating approximate labels automatically, for example from rules or user clicks, then training on many noisy labels instead of few perfect ones.
  • "Do LLMs need labels?" — Pretraining uses self-supervision: the next token is the label, taken from the text itself. Fine-tuning and RLHF use human-made labels and preferences.