Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are Autoencoders? What is the use of Autoencoders?
What you need to know
How it works
- Encoder — layers that shrink the input: 30 → 16 → 8.
- Bottleneck — the 8-number code.
- Decoder — layers that expand it back: 8 → 16 → 30.
- Loss — how different the output is from the input, usually mean squared error for numbers, binary cross-entropy for pixels in 0–1.
The target is the input itself, so every unlabelled example is training data.
1import torch, torch.nn as nn23class AutoEncoder(nn.Module):4 def __init__(self, n_features=30, code=8):5 super().__init__()6 self.encoder = nn.Sequential(nn.Linear(n_features, 16), nn.ReLU(), nn.Linear(16, code))7 self.decoder = nn.Sequential(nn.Linear(code, 16), nn.ReLU(), nn.Linear(16, n_features))89 def forward(self, x):10 return self.decoder(self.encoder(x))1112model = AutoEncoder()13opt = torch.optim.Adam(model.parameters(), lr=1e-3)14normal = torch.randn(5000, 30) # stand-in for scaled normal transactions1516for epoch in range(20):17 for batch in normal.split(256):18 loss = nn.functional.mse_loss(model(batch), batch) # target = input19 opt.zero_grad(); loss.backward(); opt.step()2021with torch.no_grad():22 new = torch.randn(4, 30)23 error = ((model(new) - new) ** 2).mean(dim=1) # one anomaly score per rowThe last line gives one reconstruction error per transaction. You pick a threshold on validation data, for example the 99.5th percentile of errors on normal transactions, and flag anything above it.
Why anomaly detection works
The autoencoder only saw normal data, so its 8-number code can describe normal patterns well. An unusual input, such as a fraud pattern it never saw, does not fit those patterns, so the decoder rebuilds it badly and the error is high. It learns "what normal looks like", which is useful when anomalies are rare and keep changing.
Main uses
- Anomaly detection — fraud, machine faults, network intrusions.
- Denoising — feed a noisy input, train it to output the clean one (a denoising autoencoder).
- Dimensionality reduction — a non-linear version of PCA; the bottleneck is a compact feature vector.
- Pretraining — masked autoencoders (MAE) hide 75% of image patches and learn to rebuild them; the encoder is then fine-tuned.
- Generation — a variational autoencoder (VAE) makes the code a probability distribution, so you can sample new codes and decode them into new data. Stable Diffusion uses a VAE to compress images before diffusion.
A real-life example
A card issuer sees about 1 fraud in every 1,000 transactions, and fraudsters change tactics monthly. A supervised model trained on last year's fraud labels misses new attack types.
They train an autoencoder on 10 million known-good transactions, with 30 scaled features each. On held-out normal transactions, the average reconstruction error is 0.05, and 99.5% are below 0.30. They set the alert threshold at 0.30.
A new scheme appears: many small ₹1 "verification" charges from a new merchant, followed by one large purchase. The supervised model scores it low risk because it has never seen this pattern. The autoencoder's error for these transactions is 0.9, three times the threshold, and they are sent for review. The team then labels the new fraud and adds it to the supervised model. The two models work together: one catches known patterns precisely, the other catches the unknown.
Follow-up questions to expect
- "How is an autoencoder different from PCA?" — PCA is linear. An autoencoder with non-linear activations can learn curved structure. A linear autoencoder with MSE loss learns the same subspace as PCA.
- "What stops an autoencoder from just copying the input?" — The bottleneck being smaller than the input, or other constraints: noise on the input, sparsity penalties, or masking.
- "What is a VAE?" — An autoencoder whose encoder outputs a mean and variance, trained with reconstruction loss plus a KL-divergence term that keeps codes close to a standard normal distribution, so sampling works.