Coding Hubs School of AI – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 8448811540
Back to Deep Learning Notes
Topic #285

Mini-Batch Gradient Descent

Mini-Batch Gradient Descent is the practical compromise between batch and stochastic gradient descent — and it's what virtually every real deep learning training run actually uses, whether or not the optimizer is casually called "SGD" in code.

Formula

\[ \nabla L(\mathbf{w}) \approx \frac{1}{B}\sum_{i=1}^{B} \nabla L_i(\mathbf{w}) \qquad \text{(a mini-batch of size } B\text{, e.g. 32, 64, 128, 256)} \]

Instead of one example (SGD) or the whole dataset (batch GD), each update uses a small, randomly sampled subset — a "mini-batch."

Why This Specific Compromise Wins

Batch GDSGDMini-Batch GD
Gradient noiseNoneHighModerate — averages out much of SGD's noise
Update frequencyOnce per epochOnce per exampleOnce per batch — frequent, but not wastefully so
GPU/hardware utilizationGood (large matrix ops), but rare updatesPoor — one example doesn't fill a GPU's parallel capacityExcellent — batches are sized to exploit GPU parallelism efficiently
Memory requirementMust hold the full dataset's gradient computationMinimalModerate, tunable via batch size

Mini-batches, especially at sizes like 32–256, are large enough to average out most of SGD's noisy variance and small enough to fit comfortably in GPU memory while keeping matrix operations efficiently parallelized — the sweet spot that made it the default.

Code — Connecting to PyTorch's DataLoader

import torch
from torch.utils.data import DataLoader, TensorDataset

X = torch.randn(10000, 5)
y = torch.randn(10000)
dataset = TensorDataset(X, y)
loader = DataLoader(dataset, batch_size=64, shuffle=True)   # this IS mini-batch gradient descent

w = torch.zeros(5, requires_grad=True)
optimizer = torch.optim.SGD([w], lr=0.01)

for epoch in range(5):
    for X_batch, y_batch in loader:   # each iteration = one mini-batch update
        predictions = X_batch @ w
        loss = ((predictions - y_batch) ** 2).mean()
        loss.backward()
        optimizer.step()
        optimizer.zero_grad()

Note that torch.optim.SGD is the same optimizer class used here, whether you feed it one example, a mini-batch, or the whole dataset at once — the "SGD vs mini-batch vs batch" distinction is about how you construct your DataLoader and training loop, not a different optimizer class.

Choosing a Batch Size

Batch SizeEffect
Small (e.g. 8–32)More noise (closer to SGD), more frequent updates, lower memory use
Large (e.g. 256–1024+)Smoother gradient estimate (closer to batch GD), fewer updates per epoch, higher memory use, often requires a proportionally larger learning rate

Common Mistakes

  • Choosing a batch size purely for GPU memory convenience without considering its effect on gradient noise and generalization — very large batch sizes can sometimes generalize slightly worse without other adjustments (like learning rate scaling).
  • Forgetting to shuffle the dataset before batching — without shuffling, each epoch sees mini-batches in the same fixed order, which can introduce unwanted correlation between consecutive updates.

Interview Relevance

Q: "When people say a model was trained with 'SGD,' what are they usually actually describing?" Almost always mini-batch gradient descent — using an optimizer like torch.optim.SGD with a DataLoader that yields batches of, say, 32 or 64 examples per update, not literal single-example stochastic gradient descent. The "SGD" naming is a historical holdover.

Practice Question

A dataset has 50,000 examples. With a batch size of 100, how many weight updates happen in one epoch?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →