Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the bottleneck in an autoencoder, and why is it used?
What you need to know
The shape of an autoencoder
An encoder maps the input down to a small latent code z, and a decoder maps z back up to a reconstruction. Training minimises the difference between input and reconstruction, often mean squared error. There are no labels; the input is its own target.
input 784 -> 256 -> 64 -> [bottleneck 16] -> 64 -> 256 -> 784A 28×28 image has 784 values. Squeezing it through 16 numbers is a 49-to-1 compression.
Why the bottleneck forces learning
If the middle layer had 784 units or more, the network could learn to pass the input straight through. Reconstruction error would be near zero, and the model would have learned nothing about the data. With 16 units, it cannot keep everything, so gradient descent pushes it to keep what helps reconstruction most: the main shapes and patterns that many inputs share. Random noise, which differs from example to example, is expensive to encode and gets dropped.
A useful link: a linear autoencoder with a bottleneck of size k and squared-error loss learns the same subspace as PCA with k components. Non-linear layers let it capture curved structure that PCA cannot.
Choosing the bottleneck size
- Too small: blurry reconstructions, important detail lost.
- Too large: close to copying; the code is not a meaningful summary.
- In practice: try several sizes and watch reconstruction error on validation data, and how well the code works for the downstream task.
Other ways to prevent copying exist: sparse autoencoders penalise how many units are active, and denoising autoencoders corrupt the input and ask for the clean version.
Anomaly detection with a bottleneck
Train only on normal data. The bottleneck learns to encode normal patterns well. An abnormal input does not fit those patterns, so its reconstruction error is high.
1import torch, torch.nn as nn23class AE(nn.Module):4 def __init__(self, d_in=30, d_z=4):5 super().__init__()6 self.enc = nn.Sequential(nn.Linear(d_in, 16), nn.ReLU(), nn.Linear(16, d_z))7 self.dec = nn.Sequential(nn.Linear(d_z, 16), nn.ReLU(), nn.Linear(16, d_in))8 def forward(self, x):9 return self.dec(self.enc(x))1011# after training on normal rows only:12err = ((model(x) - x) ** 2).mean(dim=1) # reconstruction error per row13flagged = err > threshold # e.g. 99th percentile on normal validation rowsThe 30 input features pass through a 4-number bottleneck. err is one score per row, and rows above the threshold are sent for review.
The VAE twist
In a variational autoencoder, the bottleneck outputs a mean and variance, and the code is sampled from that distribution. A penalty keeps these distributions close to a standard normal, so the latent space is smooth and you can sample new points from it to generate data.
A real-life example
A payments company wants to catch unusual merchant behaviour on UPI without labelled fraud cases. Each merchant-day is summarised as 30 features: transaction count, average amount, share of night-time transactions, number of unique payers, refund rate and so on. They train an autoencoder with a 4-unit bottleneck on three months of merchant-days from accounts with no complaints.
For a normal kirana store, reconstruction error is around 0.02. One day a merchant account that usually takes 40 small payments suddenly receives 900 payments of ₹1,999 from new payers between 1 and 3 a.m. The bottleneck has no way to represent that combination with patterns learned from normal stores, so the reconstruction error is 0.9, far above the threshold of 0.12. The account is flagged for review within the hour. The team retrains monthly, because normal behaviour shifts around festivals and sale days.
Follow-up questions to expect
- "How is an autoencoder different from PCA?" — PCA finds the best linear projection. An autoencoder with non-linear layers can learn curved, more complex structure; with only linear layers it recovers the same subspace as PCA.
- "What is an overcomplete autoencoder?" — One whose code is larger than the input. It needs another constraint, such as sparsity or input noise, to avoid learning the identity.
- "How do you pick the anomaly threshold?" — From the distribution of reconstruction errors on held-out normal data, for example the 99th percentile, then adjust based on how many alerts the review team can handle.