Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are tensors?
What you need to know
scalar rank 0 5vector rank 1 [1, 2, 3] shape (3,)matrix rank 2 [[1, 2], [3, 4]] shape (2, 2)rank 4 a batch of 32 RGB images shape (32, 3, 224, 224)The four properties that matter
- Shape —
imgs.shapeistorch.Size([32, 3, 224, 224]). Most bugs are shape bugs. - Dtype — the number type.
float32is the default for training.bfloat16andfloat16halve memory and speed up modern GPUs.int64is used for class labels and token IDs;int8for quantised inference. - Device — where it lives:
cpu,cuda(NVIDIA GPU) ormps(Apple silicon). Two tensors must be on the same device to be combined. - Gradient tracking — with
requires_grad=True, PyTorch records every operation on the tensor so it can compute derivatives later.
Memory, with numbers
That batch of 32 images has 32 × 3 × 224 × 224 = 4,816,896 values. At 4 bytes each (float32) that is about 19 MB. In bfloat16 it is about 9.6 MB. This arithmetic tells you whether a batch size will fit on your GPU.
Tensors doing work
1import torch23device = "cuda" if torch.cuda.is_available() else "cpu"4x = torch.randn(64, 30, device=device) # 64 rows, 30 features5W = torch.randn(30, 16, device=device, requires_grad=True) # a weight matrix67h = torch.relu(x @ W) # (64, 30) @ (30, 16) -> (64, 16)8h.mean().backward() # compute gradients9print(h.shape, W.grad.shape) # torch.Size([64, 16]) torch.Size([30, 16])One matrix multiply processes all 64 rows at once, which is what GPUs are built for. After backward(), W.grad has the same shape as W: one gradient per weight.
Broadcasting lets tensors of different shapes combine: adding a bias of shape (16,) to h of shape (64, 16) adds it to every row, without copying.
A real-life example
A card-fraud model scores UPI and card transactions. Each transaction has 30 numeric features (amount, hour, merchant category code, distance from the last transaction...). At peak, the service receives 2,000 transactions per second.
Instead of scoring one at a time, the service collects transactions for 10 milliseconds and scores them as one tensor of shape (20, 30). The model runs one matrix multiply per layer on the whole batch, on a GPU, in 2 ms. Scoring one by one would take 20 separate calls.
In production the tensors are float16 to halve memory, created on cuda, and wrapped in torch.inference_mode() so no gradient graph is recorded. During training, the same tensors are float32 or bfloat16, and the weights have requires_grad=True.
Follow-up questions to expect
- "What is the difference between a tensor and a matrix?" — A matrix is the rank-2 case. Tensors generalise to any rank.
- "What is broadcasting?" — Automatically expanding size-1 or missing dimensions so element-wise operations work between different shapes, without copying data.
- "Why bfloat16 over float16?" — bfloat16 keeps float32's exponent range, so values rarely overflow or underflow. Training in bfloat16 usually needs no loss scaling, while float16 often does.