Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #327

Why Normalization

As a network trains, every layer's input distribution keeps shifting — because the layers before it are simultaneously changing their own weights. Normalization techniques counteract this directly, keeping each layer's activations in a stable, well-behaved range throughout training.

The Problem: Shifting Activation Distributions

Consider a layer deep in a network. Every time the layers before it update their weights (which happens every single training step), the distribution of values this layer receives as input can shift — its mean, its variance, its overall scale. The layer has to constantly "re-adapt" to a moving target, which slows convergence and makes training more sensitive to weight initialization and learning rate. This phenomenon is often referred to as internal covariate shift, though the precise mechanism by which normalization helps remains an actively discussed topic in research — what's empirically uncontested is that normalization techniques reliably stabilize and speed up training in practice.

The Shared Formula Pattern

\[ \hat{x} = \frac{x-\mu}{\sqrt{\sigma^2+\epsilon}}, \qquad y = \gamma\hat{x}+\beta \]

Every normalization technique in this category follows this same two-step pattern: first, standardize to zero mean and unit variance (using \(\mu\) and \(\sigma^2\) computed over some specific set of values — this is exactly the standardization formula from Variance & Standard Deviation). Second, apply a learnable scale \(\gamma\) and shift \(\beta\), giving the network the freedom to undo the normalization if that turns out to be beneficial for a particular layer — without this learnable step, the network would be permanently locked into exactly zero-mean, unit-variance activations at every normalized layer, which isn't always ideal.

What Actually Differs Between Techniques

Every method in this category shares that exact formula — what changes is which set of values \(\mu\) and \(\sigma^2\) are computed over. This single design choice (which dimension(s) to normalize across) is the entire story distinguishing BatchNorm, LayerNorm, InstanceNorm, GroupNorm, and RMSNorm from one another, covered in the next five notes.

Diagram — The Question Each Technique Answers Differently

Data shape: (batch, channels, ...spatial) batch → channels → BatchNorm: per-channel, across the batch

BatchNorm, LayerNorm, InstanceNorm and GroupNorm each draw their normalization statistics from a different slice of the same underlying (batch, channel, spatial) data — the next four notes show exactly which slice each one uses.

Common Mistakes

  • Forgetting the learnable \(\gamma\) and \(\beta\) parameters entirely — without them, every normalized layer would be forced into exactly zero mean and unit variance permanently, removing flexibility the network might need.
  • Assuming all normalization techniques are interchangeable — the choice of which dimension to normalize over has real, architecture-specific consequences, covered precisely in the next four notes.

Interview Relevance

Q: "What's the shared idea behind BatchNorm, LayerNorm and their variants, and what's the one thing that actually differs between them?" Every one of them standardizes activations to zero mean and unit variance, then applies a learnable scale and shift to restore flexibility. What differs is purely which set of values the mean and variance are computed over — across the batch (BatchNorm), across a single example's features (LayerNorm), or various other groupings — and that single choice has significant practical consequences for which architectures each technique suits.

Practice Question

Why is the learnable scale-and-shift step (\(\gamma\), \(\beta\)) necessary, rather than just leaving every normalized layer at exactly zero mean and unit variance?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →