Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the difference between Conv1D, Conv2D and Conv3D?


What you need to know

Sliding dimensions versus channel dimension

A channel is a separate measurement at each position: red, green and blue for a photo, or 12 leads for an ECG. A filter always spans every input channel at once. What changes between Conv1D, 2D and 3D is the number of spatial axes the filter moves along.

Slides alongPyTorch input shapeTypical use
Conv1Dlength (time)(N, C, L)ECG, audio, sensor data, text embeddings
Conv2Dheight, width(N, C, H, W)Photos, X-rays, spectrograms
Conv3Ddepth, height, width(N, C, D, H, W)Video clips, CT and MRI volumes

N is the batch size and C the channels. PyTorch puts channels before the spatial axes. Keras and TensorFlow default to channels last, for example (N, H, W, C), so check the convention when you read code.

Parameter counts

Weights per layer = in_channels × out_channels × kernel volume, plus one bias per output channel:

Python
import torch.nn as nncount = lambda m: sum(p.numel() for p in m.parameters())print(count(nn.Conv1d(8, 16, kernel_size=5)))    # 8*16*5  + 16 = 656print(count(nn.Conv2d(3, 16, kernel_size=3)))    # 3*16*9  + 16 = 448print(count(nn.Conv3d(1, 16, kernel_size=3)))    # 1*16*27 + 16 = 448

The parameter counts look similar, but compute is very different. A Conv3D filter is applied at every depth position as well, so a clip of 16 frames costs about 16 times more than one frame, and each filter position does 27 multiplies instead of 9.

Which to choose

  • If the data is a signal over time with a few channels, use Conv1D.
  • If it is an image, including spectrograms that turn audio into a picture, use Conv2D.
  • If it is a volume where neighbouring slices are related, or a clip where motion matters, use Conv3D, or Conv2D per slice with a model on top if compute is tight.

A real-life example

A health-tech company builds three models:

  • Wearable ECG: 10-second recordings at 500 Hz from 12 leads give input shape (N, 12, 5000). A Conv1D stack finds heartbeat shapes such as an abnormal QRS complex wherever they occur in the recording.
  • Chest X-ray triage: one grayscale image, (N, 1, 512, 512). Conv2D.
  • Lung CT nodules: a scan is a stack of slices, for example 64 slices of 128×128 around a suspicious spot, giving (N, 1, 64, 128, 128). A nodule is a small 3D blob, and a vessel is a long tube that looks like a blob on a single slice. Only Conv3D, which sees neighbouring slices together, tells them apart reliably. Training is slow, so the team crops small 3D patches around candidates rather than processing whole scans.

Follow-up questions to expect

  • "Can you use Conv2D on audio?" — Yes, after converting it to a spectrogram, which is a 2D image of frequency against time. Many speech and sound models do this.
  • "What does a 1×1 convolution do?" — It mixes channels at each position without looking at neighbours, like a small dense layer applied per pixel. It is used to change the number of channels cheaply.
  • "Why not always use Conv3D for video?" — Cost. Many video models use 2D convolutions per frame plus a temporal model, or factorise a 3D filter into a 2D spatial and a 1D temporal part.