CNN interview questions covering convolution mechanics, pooling, receptive fields, and the key architectural innovations — with fully explained answers.
Q1. Why are CNNs better suited than plain MLPs for image data?
A plain MLP treats every pixel as an independent input feature, ignoring spatial structure entirely, and requires a separate weight for every pixel-to-neuron connection — for a realistically-sized image, this produces an enormous, spatially-blind parameter count. CNNs instead use small, shared kernels that slide across the image, exploiting two facts about images: nearby pixels are related (locality), and a useful pattern (an edge, a texture) can appear anywhere in the image (translation invariance via weight sharing) — this makes CNNs dramatically more parameter-efficient and better suited to visual data.
Q2. Explain what a convolution operation does, with an example.
A convolution slides a small kernel (e.g. 3×3) across the input, computing an element-wise multiply-and-sum at each position to produce one output value — repeated across the whole input, this produces a feature map. For example, a specific 3×3 kernel with values that detect vertical edges will produce high activation values exactly where a vertical edge exists in the input image, and near-zero values elsewhere — the network learns the actual kernel values during training, rather than these being hand-designed.
Q3. What is the purpose of pooling layers?
Pooling (max or average) downsamples a feature map, reducing its spatial dimensions while retaining the most important information. This serves two purposes: it reduces computational cost and parameter count in subsequent layers, and it provides a degree of translation invariance — a small shift in where a feature appears in the input has less effect on the pooled output than on the raw feature map.
Q4. How do you compute the output size of a convolution layer?
\[ O = \left\lfloor \frac{I - K + 2P}{S} \right\rfloor + 1 \]where \(I\) is input size, \(K\) is kernel size, \(P\) is padding, and \(S\) is stride. For example, a 32×32 input with a 3×3 kernel, padding 1, and stride 1: \(O = \lfloor(32 - 3 + 2)/1\rfloor + 1 = 32\) — with padding 1 and stride 1, a 3×3 kernel preserves spatial dimensions exactly, which is why "same" padding is common in practice.
Q5. What is the receptive field, and why does it matter?
The receptive field is the region of the original input image that a given neuron's output is actually influenced by. Early layers have small receptive fields (seeing only local patterns like edges); as you go deeper, each layer's receptive field grows, allowing later layers to represent increasingly global, complex patterns (parts of objects, then whole objects). Understanding receptive field size matters for tasks like segmentation, where the network needs a receptive field large enough to capture the full context needed to correctly classify each pixel.
Q6. Why use 1×1 convolutions? What do they actually accomplish?
A 1×1 convolution doesn't look at any spatial neighborhood — it operates purely across the channel dimension at each individual spatial location. Its main practical uses: reducing the number of channels (dimensionality reduction) before a more expensive larger convolution, as used extensively in Inception-style architectures; and adding extra non-linearity/mixing across channels cheaply, without changing spatial dimensions.
Q7. What problem does ResNet's skip connection solve?
As networks get very deep, they can suffer from the degradation problem — accuracy gets worse with added depth, even on the training set, which isn't explained by overfitting; it's an optimization difficulty. ResNet's skip (residual) connections let each block learn a residual function \(F(x)\) added to the original input (\(y = F(x) + x\)) rather than the full desired mapping directly — if the optimal mapping is close to identity, the residual is easy to push toward zero, and the skip connection also gives gradients a direct, unobstructed path back through the network, substantially easing optimization for very deep architectures.
Q8. What's the difference between max pooling and average pooling?
Max pooling takes the maximum value within each pooling window, which tends to preserve the strongest, most salient activations (useful for detecting whether a specific feature is present anywhere in a region). Average pooling takes the mean, producing a smoother, less spiky downsampled representation. Max pooling is more common in intermediate layers for feature detection; average pooling (specifically global average pooling) is commonly used just before the final classification layer, as a lighter-weight alternative to flattening plus a large fully-connected layer.
Q9. How does batch normalization help CNN training?
Batch normalization normalizes each layer's activations (per mini-batch) to have roughly zero mean and unit variance before scaling and shifting them with learnable parameters. This stabilizes and speeds up training by keeping activation distributions consistent across layers and training steps, allows higher learning rates to be used safely, and provides a mild regularizing effect from the noise introduced by per-batch statistics — all of which is why it became a near-standard component in CNN architectures after its introduction.