Course Content
Deep Learning Essentials
13 sections · 61 lessons
What are the components of a neural network?
What you need to know
| Component | What it does | Example |
|---|---|---|
| Input layer | Receives the features | 30 transaction features, or a 3×224×224 image |
| Hidden layers | Learn intermediate representations | Linear, Conv2d, attention blocks |
| Output layer | Produces the prediction, sized for the task | 1 logit (binary), 12 logits (12 classes), 1 number (regression) |
| Weights and biases | The learned parameters | About 11 million in ResNet-18 |
| Activation functions | Add non-linearity | ReLU, GELU, sigmoid, tanh |
| Loss function | Measures error | Cross-entropy, MSE |
| Optimiser | Updates weights using gradients | SGD with momentum, Adam, AdamW |
| Hyperparameters | Settings you choose, not learned | Learning rate, batch size, epochs, layer sizes |
| Regularisation and helpers | Stabilise training and reduce overfitting | Batch norm, layer norm, dropout, skip connections |
Where each component appears in code
Python
1import torch, torch.nn as nn23model = nn.Sequential( # architecture4 nn.Linear(30, 64), # hidden layer: weights + biases5 nn.BatchNorm1d(64), # normalisation6 nn.ReLU(), # activation7 nn.Dropout(0.3), # regularisation8 nn.Linear(64, 1), # output layer: one logit9)10loss_fn = nn.BCEWithLogitsLoss() # loss11optimizer = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01) # optimiser12batch_size, epochs = 128, 30 # hyperparametersThe input layer has no code of its own: it is simply the 30-feature tensor you pass in. The output layer's size and the loss are chosen together, based on the task.
A useful way to say it in an interview
"Architecture decides what functions the network can represent. Parameters decide which one it currently represents. The loss and optimiser decide how it moves towards a better one."
A real-life example
A crop-disease team lists the components of their model for a design review:
- Input: a 3×224×224 leaf photo, normalised to ImageNet mean and standard deviation.
- Hidden layers: a pretrained EfficientNet-B0 backbone of conv blocks with batch norm and SiLU activations, producing 1,280 features.
- Output layer:
Dropout(0.2)thenLinear(1280, 8)for 8 disease classes. - Parameters: about 4 million weights and biases, most of them pretrained.
- Loss: cross-entropy with label smoothing 0.1.
- Optimiser: AdamW, learning rate 3e-4 with cosine decay, weight decay 0.05.
- Hyperparameters: batch size 64, 25 epochs, early stopping with patience 5.
When validation accuracy stalled at 84%, listing the components made it easy to see what to try: a lower learning rate for the backbone, stronger augmentation, and a higher class weight for the rarest disease. Accuracy reached 91%.
Follow-up questions to expect
- "Is the optimiser part of the model?" — Not of the saved model used at inference. It is part of training. You save its state only to resume training.
- "Which components have learnable parameters?" — Linear and conv layers, normalisation layers (scale and shift), and embeddings. Activations, pooling and dropout have none.
- "How do you choose the output layer?" — From the task: one logit with sigmoid for binary, K logits with softmax for one of K classes, K sigmoids for multi-label, a linear unit for regression.