Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the difference between a CNN and an RNN, and when should you use each?


Where each architecture shares its weightsCNN• Same filter at every position• Finds local patterns anywhere• All positions computed at once• X-rays, product photos, video framesRNN, LSTM, GRU• Same weights at every time step• Hidden state carries history• Step t waits for step t − 1• Sales series, sensor streams, speech
A CNN shares weights across space, an RNN across time; for long text, transformers replaced both.

What you need to know

What a CNN computes

A convolutional layer holds small filters, for example 3×3. Each filter slides across the input and outputs, at every position, how strongly that patch matches its pattern. The same filter weights are reused at every position (weight sharing), so a pattern learned in one place is detected everywhere. Stacking layers builds up from edges to textures to object parts. All positions are computed at once, so CNNs run very fast on GPUs.

What an RNN computes

A recurrent neural network processes a sequence x1, x2, ..., xT in order. At each step it combines the new input with a hidden state h that summarises everything seen so far:

Text
h_t = tanh(W_x · x_t + W_h · h_(t−1) + b)

The same weights are reused at every time step (weight sharing across time). Because h_t depends on h_(t−1), step 100 cannot start until step 99 finishes, so RNNs are slow on long sequences. Plain RNNs also suffer from vanishing gradients over long spans; LSTM and GRU add gates that decide what to keep and forget, which lets them remember over hundreds of steps.

Shapes in PyTorch

Python
import torch, torch.nn as nnconv = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=3, padding=1)xray = torch.randn(8, 1, 224, 224)            # 8 grayscale imagesprint(conv(xray).shape)                       # torch.Size([8, 16, 224, 224])lstm = nn.LSTM(input_size=5, hidden_size=64, batch_first=True)sales = torch.randn(32, 90, 5)                # 32 stores, 90 days, 5 featuresout, (h, c) = lstm(sales)print(out.shape, h.shape)                     # [32, 90, 64]  [1, 32, 64]

The CNN keeps height and width and adds 16 feature maps. The LSTM returns a hidden vector for every day (out) and the final summary (h), which a linear layer turns into a forecast.

Choosing between them

DataGood choiceWhy
Photos, X-rays, satellite tilesCNN (or a vision transformer with pretraining)Local spatial patterns, position-independent
Short to medium time seriesLSTM or GRU, or a 1D CNNOrder matters, modest length
Text, long documents, long sequencesTransformerParallel training, direct long-range links
Video2D CNN per frame plus a temporal model, or a 3D CNNSpace and time both matter

The lines are not strict. 1D CNNs work well on sequences where local patterns matter, and vision transformers compete with CNNs on images when there is enough data or a good pretrained model.

A real-life example

A hospital network wants two models. The first triages chest X-rays, flagging likely pneumonia for a radiologist to read first. A pneumonia opacity is a local texture that can appear anywhere in the lung, which is exactly what shared convolutional filters detect, so the team fine-tunes a pretrained CNN.

The second forecasts daily demand for oxygen cylinders at each hospital, using the last 90 days of usage, admissions and a festival or monsoon flag. Here the answer depends on order and trend: a rise over the last two weeks means something different from a single spike. The team uses an LSTM over the 90-day window. They also try a small transformer, which forecasts similarly on this short window, so they keep the LSTM because it is cheaper to run.

Follow-up questions to expect

  • "Why have transformers replaced RNNs for text?" — RNNs must process tokens one after another, which is slow, and their memory fades over long spans. Attention lets every token look at every other token directly, and all positions train in parallel.
  • "Can a CNN handle sequences?" — Yes. A 1D convolution over time finds local patterns, and stacking dilated convolutions gives a long receptive field. These temporal CNNs are often strong baselines for time series.
  • "What is the difference between an LSTM and a GRU?" — A GRU merges the LSTM's forget and input gates into one update gate and has no separate cell state, so it has fewer parameters and often performs similarly.