Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #361

Average Pooling

Average pooling summarizes each local window by its mean value, rather than max pooling's single largest value — a gentler downsampling that preserves overall regional signal strength instead of just its peak.

Formula

\[ \text{AvgPool}(x)_{i,j} = \frac{1}{|\text{window}|}\sum_{(m,n)\in\text{window}} x_{i+m,j+n} \]

Numerical Example

Using the same 4×4 feature map as Max Pooling's worked example, with a 2×2 window, stride 2:

\[ \text{Top-left: }\frac{1+3+5+6}{4}=3.75, \qquad \text{Top-right: }\frac{2+4+1+2}{4}=2.25 \] \[ \text{Bottom-left: }\frac{3+2+1+4}{4}=2.5, \qquad \text{Bottom-right: }\frac{8+1+2+3}{4}=3.5 \] \[ \text{Result: }\begin{bmatrix}3.75 & 2.25\\2.5 & 3.5\end{bmatrix} \]

Compare directly to max pooling's result on the identical input, \(\begin{bmatrix}6&4\\4&8\end{bmatrix}\) — average pooling produces noticeably smaller, smoother values, since it's diluting each region's peak activation with its other, weaker values.

Max vs Average Pooling — When Each Is Preferred

Max PoolingAverage Pooling
PreservesThe strongest single activation in each regionThe overall average signal across each region
Typical useFeature detection tasks (most classification CNNs)Smoother downsampling; final-layer summarization (see Global Average Pooling)
Sensitivity to noiseCan be more sensitive to a single unusually large (possibly noisy) activationMore robust to a single outlier value, since it's diluted by averaging

Code

import torch
import torch.nn as nn

x = torch.tensor([[[[1.,3.,2.,4.],
                     [5.,6.,1.,2.],
                     [3.,2.,8.,1.],
                     [1.,4.,2.,3.]]]])

avg_pool = nn.AvgPool2d(kernel_size=2, stride=2)
print(avg_pool(x))
# tensor([[[[3.7500, 2.2500],
#           [2.5000, 3.5000]]]]) -- matches the manual calculation

Common Mistakes

  • Assuming average pooling is a strictly "safer" default choice than max pooling for every task — the right choice genuinely depends on the task; max pooling remains far more common for standard image classification, specifically because preserving peak activations tends to help feature detection.

Interview Relevance

Q: "When might average pooling be preferable to max pooling?" When the overall regional signal (not just its peak) carries useful information, or when you want smoother, less noise-sensitive downsampling — average pooling is also the standard choice specifically at the very end of many modern CNN architectures (as Global Average Pooling, next note), where summarizing an entire feature map's overall activation level, rather than its single peak, is the goal.

Practice Question

Apply 2×2 average pooling (stride 2) to the feature map \(\begin{bmatrix}2&1&5&3\\4&0&2&6\end{bmatrix}\), and compare the result to what max pooling gave for the same input in the previous note's practice question.

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →