Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are the supervised and unsupervised learning algorithms in Deep Learning?
What you need to know
The difference is in the training signal
- Supervised — each example has a target: an image with its label "late blight", a transaction with "fraud / not fraud". The loss compares the prediction with that target.
- Unsupervised — examples have no targets. The model learns by reconstructing its input, modelling the data distribution, or grouping similar examples.
- Self-supervised — a special kind of unsupervised learning. A target is made automatically from the data: hide a word and predict it, or hide part of an image and reconstruct it. The training then looks supervised, but no human labelled anything.
The main models in each group
| Group | Model | Typical use |
|---|---|---|
| Supervised | MLP (feedforward network) | Tabular prediction |
| Supervised | CNN | Image classification, detection |
| Supervised | RNN, LSTM, GRU | Older sequence models, small time series |
| Supervised | Transformer | Text, vision, speech, multimodal |
| Unsupervised | Autoencoder, VAE | Compression, denoising, anomaly detection |
| Unsupervised | GAN, diffusion model | Generating images, audio, synthetic data |
| Unsupervised | RBM, deep belief network, self-organising map | Historical; rarely used now |
| Self-supervised | Masked language modelling (BERT), next-token prediction (GPT) | Pretraining language models |
| Self-supervised | Contrastive learning (SimCLR, CLIP) | Pretraining image and image-text encoders |
Note that one architecture can be trained either way. A transformer is supervised when you fine-tune it on labelled sentiment data, and self-supervised when you pretrain it on raw text.
Two more types interviewers may mention
- Semi-supervised — a few labels plus many unlabelled examples. For example, train on labelled data, predict on unlabelled data, and add confident predictions as extra labels (pseudo-labelling).
- Reinforcement learning — learns from rewards, not labels. Deep RL is used in game playing, robotics, and in RLHF for aligning LLMs.
A real-life example
A crop-advisory company has 2 million unlabelled leaf photos uploaded by farmers, and only 10,000 photos labelled by agronomists.
- Self-supervised pretraining — they train a vision encoder on the 2 million photos with a contrastive objective: two random crops of the same photo should get similar embeddings, crops of different photos should not. No labels needed.
- Supervised fine-tuning — they add a classification head and train on the 10,000 labelled photos.
- Unsupervised monitoring — an autoencoder trained on normal uploads flags photos with high reconstruction error, such as blurry images or photos that are not leaves, before they reach the classifier.
Compared with training on the 10,000 labelled photos from scratch, the pretrained model makes about half as many errors. Most of the knowledge came from data nobody labelled.
Follow-up questions to expect
- "Is next-token prediction supervised or unsupervised?" — Self-supervised. It uses a supervised-style loss, but the target (the next token) comes from the text itself.
- "Is K-means a deep learning algorithm?" — No, it is classical unsupervised learning. But K-means is often run on embeddings produced by a deep network.
- "When would you choose unsupervised learning?" — When labels are expensive or impossible, as in fraud or equipment failure where new attack types appear, or when the goal is to find structure rather than predict a known target.