Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are tensors?


What you need to know

Text
scalar   rank 0   5vector   rank 1   [1, 2, 3]                      shape (3,)matrix   rank 2   [[1, 2], [3, 4]]               shape (2, 2)rank 4            a batch of 32 RGB images       shape (32, 3, 224, 224)

The four properties that matter

  • Shape — imgs.shape is torch.Size([32, 3, 224, 224]). Most bugs are shape bugs.
  • Dtype — the number type. float32 is the default for training. bfloat16 and float16 halve memory and speed up modern GPUs. int64 is used for class labels and token IDs; int8 for quantised inference.
  • Device — where it lives: cpu, cuda (NVIDIA GPU) or mps (Apple silicon). Two tensors must be on the same device to be combined.
  • Gradient tracking — with requires_grad=True, PyTorch records every operation on the tensor so it can compute derivatives later.

Memory, with numbers

That batch of 32 images has 32 × 3 × 224 × 224 = 4,816,896 values. At 4 bytes each (float32) that is about 19 MB. In bfloat16 it is about 9.6 MB. This arithmetic tells you whether a batch size will fit on your GPU.

Tensors doing work

Python
import torchdevice = "cuda" if torch.cuda.is_available() else "cpu"x = torch.randn(64, 30, device=device)                       # 64 rows, 30 featuresW = torch.randn(30, 16, device=device, requires_grad=True)   # a weight matrixh = torch.relu(x @ W)          # (64, 30) @ (30, 16) -> (64, 16)h.mean().backward()            # compute gradientsprint(h.shape, W.grad.shape)   # torch.Size([64, 16]) torch.Size([30, 16])

One matrix multiply processes all 64 rows at once, which is what GPUs are built for. After backward(), W.grad has the same shape as W: one gradient per weight.

Broadcasting lets tensors of different shapes combine: adding a bias of shape (16,) to h of shape (64, 16) adds it to every row, without copying.

A real-life example

A card-fraud model scores UPI and card transactions. Each transaction has 30 numeric features (amount, hour, merchant category code, distance from the last transaction...). At peak, the service receives 2,000 transactions per second.

Instead of scoring one at a time, the service collects transactions for 10 milliseconds and scores them as one tensor of shape (20, 30). The model runs one matrix multiply per layer on the whole batch, on a GPU, in 2 ms. Scoring one by one would take 20 separate calls.

In production the tensors are float16 to halve memory, created on cuda, and wrapped in torch.inference_mode() so no gradient graph is recorded. During training, the same tensors are float32 or bfloat16, and the weights have requires_grad=True.

Follow-up questions to expect

  • "What is the difference between a tensor and a matrix?" — A matrix is the rank-2 case. Tensors generalise to any rank.
  • "What is broadcasting?" — Automatically expanding size-1 or missing dimensions so element-wise operations work between different shapes, without copying data.
  • "Why bfloat16 over float16?" — bfloat16 keeps float32's exponent range, so values rarely overflow or underflow. Training in bfloat16 usually needs no loss scaling, while float16 often does.