Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #355

Stride

Stride controls how far the kernel moves between each application — a stride of 1 slides it one pixel at a time; a larger stride skips positions, producing a smaller, more downsampled output.

Formula for Output Size

\[ \text{Output size} = \left\lfloor\frac{W-K}{S}\right\rfloor + 1 \]

\(W\) is the input's width (or height), \(K\) is the kernel size, \(S\) is the stride.

Numerical Example

An input of width 7, kernel size 3, stride 1: \(\lfloor\frac{7-3}{1}\rfloor+1 = 5\) — the kernel visits 5 positions. With stride 2 instead: \(\lfloor\frac{7-3}{2}\rfloor+1 = 3\) — skipping every other position, producing a smaller output.

Diagram

Stride 1 every position visited Stride 2 positions skipped

A larger stride visits fewer positions, producing a smaller output feature map — a form of built-in downsampling.

Why Increase Stride

Effect of Larger StrideConsequence
Smaller output feature mapLess computation and memory for subsequent layers
Larger effective receptive field growth per layerEach subsequent layer's neurons "see" a proportionally larger region of the original input, faster
Coarser spatial resolutionSome fine-grained spatial detail is lost — a tradeoff against the computational savings

Code

import torch
import torch.nn as nn

x = torch.randn(1, 3, 32, 32)

conv_stride1 = nn.Conv2d(3, 16, kernel_size=3, stride=1)
print(conv_stride1(x).shape)   # torch.Size([1, 16, 30, 30])

conv_stride2 = nn.Conv2d(3, 16, kernel_size=3, stride=2)
print(conv_stride2(x).shape)   # torch.Size([1, 16, 15, 15]) -- noticeably smaller output

Common Mistakes

  • Forgetting stride affects output size, and hard-coding a fully-connected layer's expected input size without recomputing it after changing stride — a very common source of shape-mismatch errors when experimenting with an architecture.
  • Assuming stride is only relevant to convolution — pooling operations (covered later in this category) apply the exact same stride concept.

Interview Relevance

Q: "What's the effect of increasing a convolutional layer's stride from 1 to 2?" The kernel skips every other position, producing an output feature map roughly half the spatial size in each dimension — reducing computation and memory for subsequent layers, at the cost of some spatial resolution/detail. It's a common technique for downsampling within convolutional layers themselves, as an alternative or complement to separate pooling layers.

Practice Question

For an input of width 10, kernel size 3, and stride 2, compute the resulting output width using the formula above.

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →