Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

List the commonly used data structures in deep learning


What you need to know

Tensors, by number of dimensions

RankNameTypical useExample shape
0ScalarLoss value, learning rate()
1VectorBias, one embedding(128,)
2MatrixWeight matrix, batch of tabular rows(64, 30)
33D tensorBatch of token sequences after embedding(batch, seq_len, dim)
44D tensorBatch of images (PyTorch order)(32, 3, 224, 224)
55D tensorBatch of video clips(batch, C, frames, H, W)

The order of dimensions is a convention. PyTorch uses channels first (N, C, H, W). TensorFlow and most image libraries use channels last (N, H, W, C). A shape mix-up here is one of the most common bugs.

Python
import torchimgs = torch.rand(32, 3, 224, 224)            # 32 RGB images, PyTorch orderprint(imgs.permute(0, 2, 3, 1).shape)          # torch.Size([32, 224, 224, 3]) channels lastemb = torch.nn.Embedding(num_embeddings=10_000, embedding_dim=128)ids = torch.tensor([[5, 42, 7]])               # one sentence of 3 token IDsprint(emb(ids).shape)                          # torch.Size([1, 3, 128])

permute reorders dimensions without copying data. The embedding layer is a lookup table of shape (10000, 128): each token ID selects one row.

The other structures

  • Computational graph — a record of operations used to compute gradients (next lessons).
  • Embedding table — a matrix where each row is the learned vector for one ID: a word, a product, a user.
  • Sparse tensors — store only non-zero values. Useful for very sparse inputs such as bag-of-words or adjacency matrices of graphs.
  • Dataset and DataLoader — Dataset returns one example by index; DataLoader batches, shuffles and loads in parallel worker processes.
  • State dict — an ordered dictionary from parameter names to tensors. It is what torch.save(model.state_dict(), ...) writes.
Python
from torch.utils.data import TensorDataset, DataLoaderds = TensorDataset(torch.randn(1000, 30), torch.randint(0, 2, (1000,)))loader = DataLoader(ds, batch_size=64, shuffle=True)xb, yb = next(iter(loader))print(xb.shape, yb.shape, len(loader))   # torch.Size([64, 30]) torch.Size([64]) 16

1,000 rows in batches of 64 gives 16 batches per epoch (the last one has 40 rows).

A real-life example

A voice team builds a model to recognise 12 spoken commands. Their data moves through these structures:

  1. Each 1-second clip at 16 kHz is a 1D tensor of 16,000 samples.
  2. It is converted to a log-mel spectrogram: a 2D tensor of 40 frequency bands × 101 time frames.
  3. The DataLoader stacks 64 of them with a channel dimension: a 4D tensor (64, 1, 40, 101), so a normal 2D CNN can process it like an image.
  4. Labels are a 1D tensor of 64 integers from 0 to 11.
  5. The model's weights live in a state dict, about 250,000 numbers across 20 named tensors.

The first week's bug: the spectrogram code produced (64, 101, 40, 1), channels last. The model ran without error because the numbers still fit a conv layer, but it convolved over the wrong axes and accuracy stuck at 40%. Printing x.shape at each step found it in five minutes.

Follow-up questions to expect

  • "How is a tensor different from a NumPy array?" — Similar API, but a tensor can live on a GPU and can track operations for automatic differentiation.
  • "Why batch the data?" — GPUs are fast at large parallel matrix operations, and averaging the gradient over a batch gives a more stable update than one example.
  • "What does num_workers in DataLoader do?" — Loads and preprocesses batches in parallel processes, so the GPU does not wait for data.