Optimization interview questions covering optimizers, learning rate scheduling, regularization, and how to read training curves — with fully explained answers.
Q1. What's the difference between SGD, Momentum, and Adam?
Plain SGD updates weights using only the current mini-batch's gradient, which can be noisy and slow to navigate ravine-shaped loss surfaces. Momentum accumulates a moving average of past gradients, smoothing updates and accelerating progress in consistent directions while damping oscillation. Adam goes further, combining momentum with an adaptive, per-parameter learning rate based on a running estimate of squared gradients (similar to RMSProp) — this typically gives faster, more robust convergence with less manual learning-rate tuning, which is why it's a common default choice in practice.
Q2. Why does Adam combine momentum and RMSProp-style adaptive learning rates?
Momentum alone helps navigate consistent-direction gradients faster but uses the same learning rate for every parameter. RMSProp-style adaptation gives each parameter its own effective learning rate based on the recent magnitude of its gradients, which helps significantly when different parameters have very different gradient scales. Adam combines both benefits: momentum for smoother, more directed progress, and per-parameter adaptive scaling for handling parameters with very different gradient behaviors — the combination is generally more robust than either technique alone.
Q3. What problem does learning rate warmup solve?
At the very start of training, weights are randomly initialized and gradients can be large and somewhat unreliable — using a full, large learning rate immediately can cause unstable, erratic early updates. Warmup starts with a small learning rate and gradually increases it over the first portion of training, letting the model stabilize before using the full learning rate — particularly important for Transformer training, where warmup is close to standard practice.
Q4. Explain cosine annealing and why it might help training.
\[ \eta_t = \eta_{min} + \frac{1}{2}(\eta_{max} - \eta_{min})\left(1 + \cos\left(\frac{t}{T}\pi\right)\right) \]Cosine annealing smoothly decreases the learning rate following a cosine curve from a maximum value down to a minimum over the course of training. The smooth, gradual decrease (rather than the abrupt jumps of step decay) tends to allow the model to make large exploratory steps early in training and progressively finer, more precise adjustments as it approaches convergence, often producing better final performance than a fixed learning rate or simple step decay.
Q5. What is gradient clipping, and when is it necessary?
Gradient clipping caps the magnitude of gradients before applying the weight update — either clipping each gradient value individually, or rescaling the entire gradient vector if its overall norm exceeds a threshold. It's necessary when training is prone to exploding gradients — very deep networks, RNNs processing long sequences, or generally unstable training — where an occasional very large gradient could otherwise cause a destructively large weight update that derails training entirely.
Q6. Why might a very high learning rate cause loss to diverge (become NaN)?
A learning rate that's too high causes weight updates that overshoot the loss surface's local structure — instead of moving toward a lower-loss region, the update can land somewhere with even higher loss, and this can compound over successive steps, with the loss and gradient magnitudes growing without bound. Once values grow large enough, floating-point overflow produces infinity, and further arithmetic involving infinity produces NaN, at which point training has effectively broken and needs to be restarted with a smaller learning rate (and possibly gradient clipping).
Q7. What's the difference between L1 and L2 regularization, and how do their effects on weights differ?
Both add a penalty term to the loss based on the weights' magnitude, discouraging large weights. L1 (\(\sum|w_i|\)) tends to push many individual weights exactly to zero, producing sparse models — effectively performing a form of automatic feature selection. L2 (\(\sum w_i^2\)) penalizes large weights more heavily (due to the squaring) but tends to shrink weights toward small values without driving them exactly to zero, producing smoother, more evenly distributed weight magnitudes.
Q8. Why is weight decay in AdamW different from L2 regularization used with plain SGD?
With plain SGD, adding an L2 penalty to the loss function is mathematically equivalent to directly subtracting a fraction of the weight from itself each step (decoupled weight decay). But with Adam's adaptive per-parameter learning rates, adding an L2 penalty to the loss gets tangled up with the adaptive scaling in a way that makes it behave inconsistently as regularization. AdamW fixes this by applying weight decay directly to the weights, decoupled from the adaptive gradient computation, restoring weight decay's intended, consistent regularizing effect — which is why AdamW is generally preferred over Adam with a naive L2 penalty.
Q9. How do you diagnose whether a model is underfitting vs overfitting from its loss curves?
Underfitting: both training and validation loss remain high and don't improve much, or training loss itself fails to decrease adequately — the model isn't capturing the underlying pattern even in the data it's trained on. Overfitting: training loss keeps decreasing while validation loss plateaus and then starts increasing — a growing gap between the two curves is the classic signature, indicating the model is increasingly fitting training-specific noise rather than the genuine generalizable pattern.