Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #297

Warmup Learning Rate

Learning rate warmup does the opposite of every schedule covered so far — instead of starting high and decreasing, it starts near zero and gradually increases to the target learning rate over the first several steps, before any decay schedule takes over.

Formula (Linear Warmup)

\[ \eta_t = \eta_{\text{target}} \cdot \frac{t}{T_{\text{warmup}}} \qquad \text{for } t < T_{\text{warmup}} \]

\(T_{\text{warmup}}\) is the number of warmup steps. Once \(t\) reaches \(T_{\text{warmup}}\), the learning rate switches over to whatever main schedule follows (step decay, cosine annealing, or a constant rate).

Why Start Small and Ramp Up?

At the very start of training, a network's weights are randomly initialized and its internal statistics (activation magnitudes, gradient magnitudes) are often unstable and not yet representative of what they'll be once training settles in. A large learning rate applied immediately can cause a very large, poorly-informed first update — sometimes destabilizing training badly enough that it never recovers. Warmup lets the optimizer take small, cautious steps while the network's internal statistics stabilize, before committing to the full-sized steps the target learning rate provides.

Graph — Warmup Combined with Decay

warmup decay

A short rising warmup phase, followed by any of the decay schedules from earlier notes.

Especially Important for Transformers and Adam-Family Optimizers

Warmup has proven particularly important — sometimes essential, not just a minor tweak — for training Transformer architectures with Adam/AdamW, where the adaptive per-parameter scaling (from Adam Optimizer) can produce especially large, poorly-calibrated early updates before its moving averages (\(m_t\), \(v_t\)) have accumulated enough steps to be reliable — closely related to the bias-correction issue Adam itself addresses, but not fully solved by bias correction alone in practice.

Code

import torch.optim as optim
from torch.optim.lr_scheduler import LambdaLR

optimizer = optim.AdamW(model.parameters(), lr=0.001)
warmup_steps = 1000

def lr_lambda(step):
    if step < warmup_steps:
        return step / warmup_steps   # linear ramp from 0 to 1 (multiplied by base lr)
    return 1.0   # after warmup, hold at the target rate (or hand off to another scheduler)

scheduler = LambdaLR(optimizer, lr_lambda=lr_lambda)

Common Mistakes

  • Skipping warmup for large-batch or Transformer-based training and attributing the resulting instability to something else (bad data, bad architecture) — always check whether adding a short warmup phase resolves early-training instability before assuming a deeper problem.
  • Using an excessively long warmup phase — this wastes training budget on unnecessarily small steps once the network's statistics have already stabilized.

Interview Relevance

Q: "Why do many Transformer training recipes use a learning rate warmup phase?" At the very start of training, a randomly initialized network's activation and gradient statistics are unstable, and adaptive optimizers like Adam haven't yet accumulated enough steps for their moving averages to be reliable. A large learning rate applied immediately risks a destabilizing first update; warmup ramps the rate up gradually, letting the network and optimizer's internal statistics stabilize before full-sized updates begin.

Practice Question

With a target learning rate of 0.001 and a 500-step linear warmup, what is the learning rate at step 100?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →