Course Content
Deep Learning Essentials
13 sections · 61 lessons
List the commonly used data structures in deep learning
What you need to know
Tensors, by number of dimensions
| Rank | Name | Typical use | Example shape |
|---|---|---|---|
| 0 | Scalar | Loss value, learning rate | () |
| 1 | Vector | Bias, one embedding | (128,) |
| 2 | Matrix | Weight matrix, batch of tabular rows | (64, 30) |
| 3 | 3D tensor | Batch of token sequences after embedding | (batch, seq_len, dim) |
| 4 | 4D tensor | Batch of images (PyTorch order) | (32, 3, 224, 224) |
| 5 | 5D tensor | Batch of video clips | (batch, C, frames, H, W) |
The order of dimensions is a convention. PyTorch uses channels first (N, C, H, W). TensorFlow and most image libraries use channels last (N, H, W, C). A shape mix-up here is one of the most common bugs.
1import torch23imgs = torch.rand(32, 3, 224, 224) # 32 RGB images, PyTorch order4print(imgs.permute(0, 2, 3, 1).shape) # torch.Size([32, 224, 224, 3]) channels last56emb = torch.nn.Embedding(num_embeddings=10_000, embedding_dim=128)7ids = torch.tensor([[5, 42, 7]]) # one sentence of 3 token IDs8print(emb(ids).shape) # torch.Size([1, 3, 128])permute reorders dimensions without copying data. The embedding layer is a lookup table of shape (10000, 128): each token ID selects one row.
The other structures
- Computational graph — a record of operations used to compute gradients (next lessons).
- Embedding table — a matrix where each row is the learned vector for one ID: a word, a product, a user.
- Sparse tensors — store only non-zero values. Useful for very sparse inputs such as bag-of-words or adjacency matrices of graphs.
- Dataset and DataLoader —
Datasetreturns one example by index;DataLoaderbatches, shuffles and loads in parallel worker processes. - State dict — an ordered dictionary from parameter names to tensors. It is what
torch.save(model.state_dict(), ...)writes.
1from torch.utils.data import TensorDataset, DataLoader23ds = TensorDataset(torch.randn(1000, 30), torch.randint(0, 2, (1000,)))4loader = DataLoader(ds, batch_size=64, shuffle=True)5xb, yb = next(iter(loader))6print(xb.shape, yb.shape, len(loader)) # torch.Size([64, 30]) torch.Size([64]) 161,000 rows in batches of 64 gives 16 batches per epoch (the last one has 40 rows).
A real-life example
A voice team builds a model to recognise 12 spoken commands. Their data moves through these structures:
- Each 1-second clip at 16 kHz is a 1D tensor of 16,000 samples.
- It is converted to a log-mel spectrogram: a 2D tensor of 40 frequency bands × 101 time frames.
- The
DataLoaderstacks 64 of them with a channel dimension: a 4D tensor(64, 1, 40, 101), so a normal 2D CNN can process it like an image. - Labels are a 1D tensor of 64 integers from 0 to 11.
- The model's weights live in a state dict, about 250,000 numbers across 20 named tensors.
The first week's bug: the spectrogram code produced (64, 101, 40, 1), channels last. The model ran without error because the numbers still fit a conv layer, but it convolved over the wrong axes and accuracy stuck at 40%. Printing x.shape at each step found it in five minutes.
Follow-up questions to expect
- "How is a tensor different from a NumPy array?" — Similar API, but a tensor can live on a GPU and can track operations for automatic differentiation.
- "Why batch the data?" — GPUs are fast at large parallel matrix operations, and averaging the gradient over a batch gives a more stable update than one example.
- "What does
num_workersin DataLoader do?" — Loads and preprocesses batches in parallel processes, so the GPU does not wait for data.