Course Content
Deep Learning Essentials
13 sections · 61 lessons
What is the learning rate in a neural network?
What you need to know
The update rule
w_new = w_old − learning_rate × ∂Loss/∂wThe gradient gives the direction and the local steepness. The learning rate decides how far to move. Because the gradient only describes the slope where you stand, a large step can land somewhere the slope is completely different.
Watch it on a simple bowl
Take the loss L = w², whose minimum is at w = 0. The gradient is 2w. Start at w = 1:
- lr = 0.1: each step does
w = w − 0.2w = 0.8w. The sequence is 1, 0.8, 0.64, 0.51... and it converges smoothly. - lr = 1.0: each step does
w = w − 2w = −w. The sequence is 1, −1, 1, −1... and it bounces forever. - lr = 1.1: each step does
w = −1.2w. The sequence is 1, −1.2, 1.44, −1.73... and it diverges. - lr = 0.001: each step does
w = 0.998w. After 100 stepswis still 0.82. That is crawling.
Real loss surfaces are not one bowl, but the same three behaviours show up in every training run.
Schedules: not one number but a curve
- Warmup: start small (for example, 0 rising to the target over the first 1–5% of steps). Early gradients are large and Adam's estimates are poor, so a big early step can wreck the weights.
- Decay: reduce the rate later so the weights settle into a minimum instead of bouncing around it. Common choices are cosine decay, step decay and
ReduceLROnPlateau, which cuts the rate when validation loss stops improving.
1opt = torch.optim.AdamW(model.parameters(), lr=3e-4, weight_decay=0.01)2sched = torch.optim.lr_scheduler.OneCycleLR(3 opt, max_lr=3e-4, total_steps=10_000, pct_start=0.05) # 5% warmup, then decay4# call sched.step() after every optimizer.step()Choosing a starting value
A learning-rate range test trains for a few hundred steps while increasing the rate exponentially, for example from 1e-7 to 1, and plots loss against rate. The loss falls, flattens, then shoots up. Pick a value roughly 10 times smaller than the point where the loss is lowest.
The right rate also depends on the situation: fine-tuning a pretrained model needs a much smaller rate than training from scratch, because big steps destroy what the model already knows.
A real-life example
A team fine-tunes a pretrained BERT-style model to classify customer-support tickets for a telecom company into 15 categories. Copying a from-scratch recipe, they use a learning rate of 0.01. Training loss falls briefly, then settles at the level of always predicting the most common class. The large steps overwrote the pretrained language knowledge within the first few hundred steps.
They drop the rate to 2e-5, a common range for BERT fine-tuning, and add 10% warmup with linear decay. Accuracy on held-out tickets climbs steadily and ends well above the first run. Nothing else changed.
Follow-up questions to expect
- "How does the learning rate interact with batch size?" — Larger batches give less noisy gradients, so they can usually take a larger rate. A common rule is to scale the rate in proportion to the batch size, with warmup.
- "Does Adam remove the need to tune the learning rate?" — No. Adam adapts step sizes relative to each other, but the global rate still decides whether training is stable.
- "What is
ReduceLROnPlateau?" — A scheduler that watches a metric such as validation loss and multiplies the rate by a factor, often 0.1, when the metric has not improved for a set number of epochs.