Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the difference between a CNN and an RNN, and when should you use each?
What you need to know
What a CNN computes
A convolutional layer holds small filters, for example 3×3. Each filter slides across the input and outputs, at every position, how strongly that patch matches its pattern. The same filter weights are reused at every position (weight sharing), so a pattern learned in one place is detected everywhere. Stacking layers builds up from edges to textures to object parts. All positions are computed at once, so CNNs run very fast on GPUs.
What an RNN computes
A recurrent neural network processes a sequence x1, x2, ..., xT in order. At each step it combines the new input with a hidden state h that summarises everything seen so far:
h_t = tanh(W_x · x_t + W_h · h_(t−1) + b)The same weights are reused at every time step (weight sharing across time). Because h_t depends on h_(t−1), step 100 cannot start until step 99 finishes, so RNNs are slow on long sequences. Plain RNNs also suffer from vanishing gradients over long spans; LSTM and GRU add gates that decide what to keep and forget, which lets them remember over hundreds of steps.
Shapes in PyTorch
1import torch, torch.nn as nn23conv = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=1)4xray = torch.randn(8, 1, 224, 224) # 8 grayscale images5print(conv(xray).shape) # torch.Size([8, 16, 224, 224])67lstm = nn.LSTM(input_size=5, hidden_size=64, batch_first=True)8sales = torch.randn(32, 90, 5) # 32 stores, 90 days, 5 features9out, (h, c) = lstm(sales)10print(out.shape, h.shape) # [32, 90, 64] [1, 32, 64]The CNN keeps height and width and adds 16 feature maps. The LSTM returns a hidden vector for every day (out) and the final summary (h), which a linear layer turns into a forecast.
Choosing between them
| Data | Good choice | Why |
|---|---|---|
| Photos, X-rays, satellite tiles | CNN (or a vision transformer with pretraining) | Local spatial patterns, position-independent |
| Short to medium time series | LSTM or GRU, or a 1D CNN | Order matters, modest length |
| Text, long documents, long sequences | Transformer | Parallel training, direct long-range links |
| Video | 2D CNN per frame plus a temporal model, or a 3D CNN | Space and time both matter |
The lines are not strict. 1D CNNs work well on sequences where local patterns matter, and vision transformers compete with CNNs on images when there is enough data or a good pretrained model.
A real-life example
A hospital network wants two models. The first triages chest X-rays, flagging likely pneumonia for a radiologist to read first. A pneumonia opacity is a local texture that can appear anywhere in the lung, which is exactly what shared convolutional filters detect, so the team fine-tunes a pretrained CNN.
The second forecasts daily demand for oxygen cylinders at each hospital, using the last 90 days of usage, admissions and a festival or monsoon flag. Here the answer depends on order and trend: a rise over the last two weeks means something different from a single spike. The team uses an LSTM over the 90-day window. They also try a small transformer, which forecasts similarly on this short window, so they keep the LSTM because it is cheaper to run.
Follow-up questions to expect
- "Why have transformers replaced RNNs for text?" — RNNs must process tokens one after another, which is slow, and their memory fades over long spans. Attention lets every token look at every other token directly, and all positions train in parallel.
- "Can a CNN handle sequences?" — Yes. A 1D convolution over time finds local patterns, and stacking dilated convolutions gives a long receptive field. These temporal CNNs are often strong baselines for time series.
- "What is the difference between an LSTM and a GRU?" — A GRU merges the LSTM's forget and input gates into one update gate and has no separate cell state, so it has fewer parameters and often performs similarly.