Coding Hubs School of AI – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 8448811540
Back to Deep Learning Notes
Topic #284

Stochastic Gradient Descent

Stochastic Gradient Descent (SGD) sits at the opposite extreme from batch gradient descent: it updates the weights using the gradient from just one randomly chosen training example at a time.

Formula

\[ \mathbf{w}_{t+1} = \mathbf{w}_t - \eta\nabla L_i(\mathbf{w}_t) \qquad \text{(a single, randomly sampled example } i\text{)} \]

Why "Stochastic"

Each individual example's gradient is a noisy, imperfect estimate of the true gradient over the full dataset — but it's an unbiased estimate (its expected value, averaged over many random examples, equals the true full-dataset gradient — see Expected Value). This randomness is the "stochastic" part, and it fundamentally changes the optimization's behavior.

The Noisy Path — A Feature, Not Just a Bug

Batch GD (smooth) SGD (noisy, zig-zag)

SGD's noisy path takes a less direct route, but that same noise can help it "jump out of" shallow local minima that trap smoother optimizers.

This noise is a real, useful property: in non-convex loss surfaces (common in deep learning, see Gradient Descent), the randomness can help SGD escape shallow local minima that smoother methods like batch gradient descent might get stuck in.

Advantages and Disadvantages

AdvantageDisadvantage
Extremely fast per-update — no waiting for a full dataset passHigh variance in the update direction; convergence path is noisy
Can escape shallow local minima due to noiseDoesn't fully exploit modern hardware's parallelism (processing one example at a time is inefficient on a GPU)
Can start updating weights before seeing the whole datasetRequires a smaller learning rate to avoid excessive oscillation

Code

import torch
import random

def sgd_step(X, y, w, lr=0.01):
    i = random.randint(0, len(X) - 1)   # pick ONE random example
    x_i, y_i = X[i], y[i]
    prediction = x_i @ w
    loss = (prediction - y_i) ** 2
    loss.backward()
    with torch.no_grad():
        w -= lr * w.grad
    w.grad.zero_()
    return w

Common Mistakes

  • Confusing "SGD" as used casually in deep learning frameworks (e.g. torch.optim.SGD) with true one-example-at-a-time SGD — in practice, "SGD" almost always refers to mini-batch gradient descent (next note) under the hood; pure single-example SGD is rarely used directly due to poor hardware utilization.
  • Using too large a learning rate with pure SGD — the high per-step variance compounds with an aggressive learning rate, often causing divergence.

Interview Relevance

Q: "What's the practical benefit of SGD's noisy gradient estimates, beyond just being faster to compute?" The noise can act as an implicit regularizer and help the optimizer escape shallow local minima or saddle points that a smoother, exact-gradient method (like batch gradient descent) might get stuck at — a genuinely useful side effect of the randomness, not just a necessary evil.

Practice Question

Why does SGD typically require a smaller learning rate than batch gradient descent to train stably?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →