Coding Now – Best AI & Full Stack Courses in Delhi NCR | 100% Placement
Limited Offer: Get 50% OFF on AI & Full Stack Courses
📞 Call Now: +91 9667708830
Back to Deep Learning Notes
Topic #379

Object Detection

Object detection extends classification with a second, simultaneous job: not just naming what's in an image, but drawing a bounding box around where each object is — and doing this for a variable number of objects per image.

The Task, Precisely

For each object detected, a model outputs a bounding box (typically \(x_1,y_1,x_2,y_2\) coordinates), a predicted class, and a confidence score. Unlike classification's single fixed-size output, detection must handle anywhere from zero to many objects per image — a fundamentally different output structure.

Two Broad Families of Approaches

FamilyApproachExamples
Two-stageFirst propose candidate regions, then classify/refine each oneR-CNN, Fast R-CNN, Faster R-CNN (covered later in this category)
One-stage (single-shot)Predict boxes and classes directly in one pass, no separate proposal stepSSD, YOLO (covered later in this category)

Two-stage approaches are generally more accurate but slower; one-stage approaches trade some accuracy for substantially faster inference, often enabling real-time detection.

Evaluation

Object detection is evaluated using exactly the metrics already covered in the Evaluation Metrics category: IoU (see IoU) determines whether a predicted box counts as a correct match, and mean Average Precision (see Mean Average Precision) summarizes overall detection quality across classes and confidence thresholds.

Code — Using a Pretrained Detector

import torchvision.models.detection as detection_models

model = detection_models.fasterrcnn_resnet50_fpn(weights='DEFAULT')
model.eval()

x = [torch.randn(3, 480, 640)]   # detection models expect a list of images
predictions = model(x)
print(predictions[0].keys())   # dict_keys(['boxes', 'labels', 'scores'])

Common Mistakes

  • Confusing object detection with plain image classification — detection requires localizing (and counting) a variable number of objects, a structurally different problem from assigning one label to the whole image.
  • Evaluating a detector with only a single IoU threshold when comparing against benchmarks that report mAP averaged across many thresholds — always match the exact metric variant when comparing numbers.

Interview Relevance

Q: "What's the fundamental difference between one-stage and two-stage object detectors?" Two-stage detectors first generate candidate object regions, then classify and refine each candidate separately — generally more accurate but slower. One-stage detectors predict boxes and classes directly across the image in a single pass, without a separate proposal step — faster, often enabling real-time detection, at some historical cost to accuracy (though this gap has narrowed considerably with modern one-stage architectures).

Practice Question

Why can't object detection be framed as a standard fixed-size classification problem, the way whole-image classification can?

Related DL Notes

Want to go beyond the notes?

Join Coding Hubs School of AI's Deep Learning course — live mentorship, real projects, and 100% placement support.

Enroll Now — Free Demo Available
💬 Talk to Advisor
1
WhatsApp

Latest from Our Blog

Insights on AI, Data Science, Full Stack & Career

View All Articles →