VGG (2014) took a strikingly simple idea to its logical extreme: use nothing but small, uniform 3×3 convolutions, stacked very deep (16 to 19 layers), proving that architectural simplicity combined with sufficient depth could match or exceed more complex designs.
The Problem It Solved
AlexNet used varied, fairly large kernel sizes (11×11, 5×5) somewhat arbitrarily chosen. VGG asked a more disciplined question: what happens if you use only the smallest practical kernel size (3×3) throughout, relying entirely on depth — rather than kernel size — to build up a large effective receptive field?
Architecture
VGG stacks blocks of two or three consecutive 3×3 convolutional layers (each followed by ReLU), interspersed with 2×2 max pooling, repeated many times, followed by large fully-connected layers. VGG-16 and VGG-19 (named for their total layer counts) are the most commonly referenced variants.
Key Innovation: Depth via Uniform Small Kernels
As established in Receptive Field, stacking three 3×3 convolutional layers achieves the same 7×7 effective receptive field as one large 7×7 kernel, but with fewer parameters and more non-linear activations along the way. VGG's entire design philosophy is built directly on this principle, applied consistently and simply throughout the whole network — depth and uniformity over cleverness in kernel size choice.
Advantages and Limitations
| Advantages | Limitations |
|---|---|
| Simple, uniform, easy-to-understand and easy-to-extend design | An enormous number of parameters (VGG-16 has roughly 138 million), mostly in its fully-connected layers |
| Strong, reliable accuracy; became a common baseline and transfer-learning starting point for years | Computationally expensive to train and to run inference with, relative to more modern, more efficient architectures |
Use Cases
Still commonly used today specifically as a pretrained feature extractor for transfer learning and as the backbone for perceptual loss functions (measuring similarity between images via VGG's intermediate feature representations) — even though it's rarely chosen for training a brand-new model from scratch given more efficient modern alternatives.
Common Mistakes
- Assuming VGG's huge parameter count means it's simply "more powerful" than more parameter-efficient modern architectures — much of that parameter count comes from its large fully-connected layers rather than genuinely useful extra representational capacity, exactly the inefficiency Global Average Pooling (see Global Average Pooling) was later developed to address.
Interview Relevance
Q: "What was VGG's key architectural insight, and what was its main drawback?" VGG showed that using only small, uniform 3×3 convolutions, stacked very deep, could match or exceed the accuracy of architectures using larger, more varied kernel sizes — relying on depth rather than kernel size for a large effective receptive field. Its main drawback is an enormous parameter count, largely concentrated in its fully-connected layers, making it computationally expensive relative to more parameter-efficient later architectures.
Practice Question
Using the receptive-field growth formula from Receptive Field, confirm that stacking two 3×3 convolutions (stride 1) achieves the same receptive field as one 5×5 convolution.