Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #388

R-CNN

R-CNN (Regions with CNN features, 2014) was the first architecture to successfully combine region proposals with CNN-based classification for object detection — establishing the two-stage detection paradigm, at the cost of being extremely slow.

The Problem It Solved

Before R-CNN, object detection relied heavily on hand-engineered features and sliding-window classifiers — exhaustively checking every possible location and scale in an image was both computationally expensive and produced comparatively weak features. R-CNN asked: what if a separate, cheaper algorithm first proposed a much smaller set of likely object regions, and only those were run through a (much more powerful) CNN?

The Three-Step Pipeline

  1. Region proposal: a separate, non-neural algorithm (Selective Search) generates roughly 2,000 candidate object regions per image.
  2. Feature extraction: each of these ~2,000 regions is individually cropped, resized, and passed separately through a CNN to extract features.
  3. Classification: each region's extracted features are classified (e.g. with an SVM) into an object class or background.

The Fatal Flaw: Extreme Slowness

Running the CNN separately for each of ~2,000 proposed regions, per image, means the same image's overlapping regions get redundantly re-processed by the CNN over and over — R-CNN could take tens of seconds per image, making it entirely impractical for any real-time or large-scale use. This single bottleneck is exactly what Fast R-CNN, covered next, was designed to eliminate.

Advantages and Limitations

AdvantagesLimitations
First successful combination of region proposals with CNN features, a major accuracy leap over prior methodsExtremely slow — ~2,000 separate CNN forward passes per image
Established the two-stage detection paradigm still influential todayTraining required multiple separate stages (CNN fine-tuning, SVM training, box regression), not end-to-end

Use Cases

R-CNN itself is purely of historical interest today — its successors (Fast R-CNN, Faster R-CNN) fixed its core inefficiency almost immediately and are what's actually used in any modern two-stage detector.

Common Mistakes

  • Assuming R-CNN's core idea (region proposals + CNN classification) was flawed — the idea itself was sound and influential; the specific implementation's redundant per-region CNN passes were the fixable bottleneck.

Interview Relevance

Q: "What made the original R-CNN so slow, and what was the key insight that fixed it?" R-CNN ran a full CNN forward pass separately for each of roughly 2,000 proposed regions per image, with massive redundant computation across overlapping regions. Fast R-CNN's key insight was to run the CNN just once per whole image, then extract per-region features from that single shared feature map — eliminating the redundant computation entirely.

Practice Question

Roughly how many separate CNN forward passes does R-CNN perform per image, and why is that number so costly?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →