Coding Hubs School of AI – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 8448811540
Back to Deep Learning Notes
Topic #350

Why CNN?

This note makes the case for CNNs concrete and quantitative — showing exactly why feeding an image into a plain fully-connected network is impractical, and precisely which two design choices in convolution fix that.

The Parameter Explosion Problem with a Plain MLP

Consider a modest 224×224 RGB image, flattened into a single vector: \(224\times224\times3 = 150{,}528\) values. A single fully-connected hidden layer with just 1,000 neurons would need a weight matrix of shape \((1000, 150{,}528)\) — over 150 million parameters, for one layer alone, on a fairly small image. This is both computationally wasteful and prone to severe overfitting, especially with limited training data.

Two Design Choices That Fix This

Design ChoiceWhat It MeansWhy It Helps
Sparse (local) connectivityEach output value depends only on a small local neighborhood of the input (the kernel's receptive field), not the entire imageDramatically fewer connections per output, matching the intuition that a pixel's meaning mostly depends on its immediate neighbors
Parameter sharingThe exact same kernel (weights) is reused at every spatial position across the imageThe number of learnable parameters depends only on the kernel size, not the image size — a small 3×3 kernel has just 9 weights (plus channels), regardless of whether the image is 32×32 or 4000×4000

Numerical Comparison

For that same 224×224×3 image, a single convolutional layer with 64 filters of size 3×3 has \(64\times(3\times3\times3+1) = 1{,}792\) parameters (each filter: 3×3 spatial × 3 input channels, plus one bias) — compare this to the 150+ million parameters a single fully-connected layer with even modestly-sized output would need. This roughly five-order-of-magnitude reduction is exactly why CNNs became practical for image tasks where MLPs were not.

Translation Invariance — A Bonus Property

Because the same kernel scans across every position, a pattern (like an edge or a specific texture) is detected wherever it appears in the image — a cat in the top-left corner activates the same filters as a cat in the bottom-right corner. Plain MLPs have no such built-in property; they'd need to separately learn to recognize a pattern at every possible position, using entirely separate weights for each.

Code — Comparing Parameter Counts Directly

import torch.nn as nn

# A single fully-connected layer processing a small flattened 32x32x3 image
fc_layer = nn.Linear(32*32*3, 1000)
print(sum(p.numel() for p in fc_layer.parameters()))   # over 3 million parameters

# A convolutional layer processing the SAME image, unflattened
conv_layer = nn.Conv2d(in_channels=3, out_channels=64, kernel_size=3, padding=1)
print(sum(p.numel() for p in conv_layer.parameters()))   # under 2,000 parameters

Common Mistakes

  • Assuming CNNs are simply "smaller" MLPs by coincidence — the parameter reduction is a direct, structural consequence of sparse connectivity and parameter sharing, not an incidental difference.
  • Overstating translation invariance as perfect — CNNs are only approximately translation-invariant in practice (pooling and certain architectural choices contribute to this), not perfectly invariant under all transformations.

Interview Relevance

Q: "What two specific properties of convolution make CNNs far more parameter-efficient than fully-connected networks for image data?" Sparse (local) connectivity — each output depends only on a small local neighborhood, not the entire input — and parameter sharing — the same small set of weights (the kernel) is reused at every spatial position, so the parameter count doesn't scale with image size at all, only with kernel size.

Practice Question

A 64×64×3 image is fed into a fully-connected layer with 500 output neurons. How many weight parameters does this single layer have? Compare this to a convolutional layer with 32 filters of size 5×5 on the same input.

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →