Deep Learning Essentials

Course Content

Deep Learning Essentials

13 sections · 61 lessons

What is the learning rate in a neural network?


Minimising w squared from w = 1 at four learning rates10.80.640.511−11−11−1.21.44−1.7310.9980.9960.994step 0step 1step 2step 3lr 0.1lr 1.0lr 1.1lr 0.001Converges, bounces forever, diverges, crawls.
The same gradient gives four different outcomes; only the step size changed.

What you need to know

The update rule

Text
w_new = w_old − learning_rate × ∂Loss/∂w

The gradient gives the direction and the local steepness. The learning rate decides how far to move. Because the gradient only describes the slope where you stand, a large step can land somewhere the slope is completely different.

Watch it on a simple bowl

Take the loss L = w², whose minimum is at w = 0. The gradient is 2w. Start at w = 1:

  • lr = 0.1: each step does w = w − 0.2w = 0.8w. The sequence is 1, 0.8, 0.64, 0.51... and it converges smoothly.
  • lr = 1.0: each step does w = w − 2w = −w. The sequence is 1, −1, 1, −1... and it bounces forever.
  • lr = 1.1: each step does w = −1.2w. The sequence is 1, −1.2, 1.44, −1.73... and it diverges.
  • lr = 0.001: each step does w = 0.998w. After 100 steps w is still 0.82. That is crawling.

Real loss surfaces are not one bowl, but the same three behaviours show up in every training run.

Schedules: not one number but a curve

  • Warmup: start small (for example, 0 rising to the target over the first 1–5% of steps). Early gradients are large and Adam's estimates are poor, so a big early step can wreck the weights.
  • Decay: reduce the rate later so the weights settle into a minimum instead of bouncing around it. Common choices are cosine decay, step decay and ReduceLROnPlateau, which cuts the rate when validation loss stops improving.
Python
opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)sched = torch.optim.lr_scheduler.OneCycleLR(    opt, max_lr=3e-4, total_steps=10_000, pct_start=0.05)  # 5% warmup, then decay# call sched.step() after every optimizer.step()

Choosing a starting value

A learning-rate range test trains for a few hundred steps while increasing the rate exponentially, for example from 1e-7 to 1, and plots loss against rate. The loss falls, flattens, then shoots up. Pick a value roughly 10 times smaller than the point where the loss is lowest.

The right rate also depends on the situation: fine-tuning a pretrained model needs a much smaller rate than training from scratch, because big steps destroy what the model already knows.

A real-life example

A team fine-tunes a pretrained BERT-style model to classify customer-support tickets for a telecom company into 15 categories. Copying a from-scratch recipe, they use a learning rate of 0.01. Training loss falls briefly, then settles at the level of always predicting the most common class. The large steps overwrote the pretrained language knowledge within the first few hundred steps.

They drop the rate to 2e-5, a common range for BERT fine-tuning, and add 10% warmup with linear decay. Accuracy on held-out tickets climbs steadily and ends well above the first run. Nothing else changed.

Follow-up questions to expect

  • "How does the learning rate interact with batch size?" — Larger batches give less noisy gradients, so they can usually take a larger rate. A common rule is to scale the rate in proportion to the batch size, with warmup.
  • "Does Adam remove the need to tune the learning rate?" — No. Adam adapts step sizes relative to each other, but the global rate still decides whether training is stable.
  • "What is ReduceLROnPlateau?" — A scheduler that watches a metric such as validation loss and multiplies the rate by a factor, often 0.1, when the metric has not improved for a set number of epochs.