Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What are the components of a neural network?


What you need to know

ComponentWhat it doesExample
Input layerReceives the features30 transaction features, or a 3×224×224 image
Hidden layersLearn intermediate representationsLinear, Conv2d, attention blocks
Output layerProduces the prediction, sized for the task1 logit (binary), 12 logits (12 classes), 1 number (regression)
Weights and biasesThe learned parametersAbout 11 million in ResNet-18
Activation functionsAdd non-linearityReLU, GELU, sigmoid, tanh
Loss functionMeasures errorCross-entropy, MSE
OptimiserUpdates weights using gradientsSGD with momentum, Adam, AdamW
HyperparametersSettings you choose, not learnedLearning rate, batch size, epochs, layer sizes
Regularisation and helpersStabilise training and reduce overfittingBatch norm, layer norm, dropout, skip connections

Where each component appears in code

Python
import torch, torch.nn as nnmodel = nn.Sequential(                    # architecture    nn.Linear(30, 64),                    # hidden layer: weights + biases    nn.BatchNorm1d(64),                   # normalisation    nn.ReLU(),                            # activation    nn.Dropout(0.3),                      # regularisation    nn.Linear(64, 1),                     # output layer: one logit)loss_fn = nn.BCEWithLogitsLoss()                                             # lossoptimizer = torch.optim.AdamW(model.parameters(), lr=1e-3, weight_decay=0.01)  # optimiserbatch_size, epochs = 128, 30                                                  # hyperparameters

The input layer has no code of its own: it is simply the 30-feature tensor you pass in. The output layer's size and the loss are chosen together, based on the task.

A useful way to say it in an interview

"Architecture decides what functions the network can represent. Parameters decide which one it currently represents. The loss and optimiser decide how it moves towards a better one."

A real-life example

A crop-disease team lists the components of their model for a design review:

  • Input: a 3×224×224 leaf photo, normalised to ImageNet mean and standard deviation.
  • Hidden layers: a pretrained EfficientNet-B0 backbone of conv blocks with batch norm and SiLU activations, producing 1,280 features.
  • Output layer: Dropout(0.2) then Linear(1280, 8) for 8 disease classes.
  • Parameters: about 4 million weights and biases, most of them pretrained.
  • Loss: cross-entropy with label smoothing 0.1.
  • Optimiser: AdamW, learning rate 3e-4 with cosine decay, weight decay 0.05.
  • Hyperparameters: batch size 64, 25 epochs, early stopping with patience 5.

When validation accuracy stalled at 84%, listing the components made it easy to see what to try: a lower learning rate for the backbone, stronger augmentation, and a higher class weight for the rarest disease. Accuracy reached 91%.

Follow-up questions to expect

  • "Is the optimiser part of the model?" — Not of the saved model used at inference. It is part of training. You save its state only to resume training.
  • "Which components have learnable parameters?" — Linear and conv layers, normalisation layers (scale and shift), and embeddings. Activations, pooling and dropout have none.
  • "How do you choose the output layer?" — From the task: one logit with sigmoid for binary, K logits with softmax for one of K classes, K sigmoids for multi-label, a linear unit for regression.