Lesson 40's sliding-window detector classified thousands of fixed-size windows and merged the results with non-maximum suppression (NMS). That works, but it is fundamentally a classification approach bolted onto a search — the network never predicts a box directly, only "face or not, at this exact window." Modern detectors instead treat localization as a regression problem: given an image, directly predict bounding box coordinates. This lesson builds the simplest possible version of that idea, then surveys how real detectors (R-CNN, YOLO, SSD) scale it up.
import numpy as np
import torch
import torch.nn as nn
import torch.nn.functional as F
import matplotlib.pyplot as plt
import matplotlib.patches as patches
The simplest possible regression problem: one circular blob per image, at an unknown location and size. Instead of a class label, the network's target is now four numbers — (cx, cy, width, height) indicating the bounding box, normalized to [0, 1] by image size.
SIZE = 32
def make_scene(rng, size=SIZE, obj_size=8):
scene = np.zeros((size, size), dtype=np.float32)
cx = rng.integers(obj_size, size - obj_size)
cy = rng.integers(obj_size, size - obj_size)
yy, xx = np.mgrid[0:size, 0:size]
scene[((xx - cx) ** 2 + (yy - cy) ** 2) <= (obj_size * 0.5) ** 2] = 1.0
scene = np.clip(scene + rng.normal(0, 0.05, scene.shape), 0, 1).astype(np.float32)
box = (cx, cy, obj_size, obj_size) # cx, cy, w, h
return scene, box
rng = np.random.default_rng(9)
N = 400
scenes, boxes = [], []
for _ in range(N):
s, b = make_scene(rng)
scenes.append(s); boxes.append(b)
scenes = np.array(scenes, dtype=np.float32)
boxes = np.array(boxes, dtype=np.float32)
split = int(0.85 * N)
Xtr, Btr = scenes[:split], boxes[:split] / SIZE
Xte, Bte = scenes[split:], boxes[split:] / SIZE
fig, axes = plt.subplots(1, 4, figsize=(9, 2.5))
for ax, im, b in zip(axes, Xtr[:4], boxes[:4]):
ax.imshow(im, cmap='gray')
ax.add_patch(patches.Rectangle((b[0] - b[2] / 2, b[1] - b[3] / 2), b[2], b[3], edgecolor='lime', facecolor='none', linewidth=2))
ax.axis('off')
plt.show()
The network is a CNN backbone (Lesson 34's pattern) followed by a 4-output regression head with a sigmoid, so every prediction lands in [0, 1] — a valid normalized box coordinate. It's trained with plain MSE loss against the true box, and evaluated with IoU (Lesson 40's intersection-over-union), the metric that actually matters for detection: how much the predicted and true boxes overlap, not how close the four numbers are in isolation.
class Detector(nn.Module):
def __init__(self):
super().__init__()
self.conv = nn.Sequential(
nn.Conv2d(1, 16, 5, padding=2), nn.ReLU(), nn.MaxPool2d(2),
nn.Conv2d(16, 32, 5, padding=2), nn.ReLU(), nn.AdaptiveMaxPool2d(1),
)
self.fc = nn.Linear(32, 4) # cx, cy, w, h, normalized
def forward(self, x):
return torch.sigmoid(self.fc(self.conv(x).flatten(1)))
def iou_batch(pred, target):
pcx, pcy, pw, ph = pred[:, 0], pred[:, 1], pred[:, 2], pred[:, 3]
tcx, tcy, tw, th = target[:, 0], target[:, 1], target[:, 2], target[:, 3]
px0, py0, px1, py1 = pcx - pw / 2, pcy - ph / 2, pcx + pw / 2, pcy + ph / 2
tx0, ty0, tx1, ty1 = tcx - tw / 2, tcy - th / 2, tcx + tw / 2, tcy + th / 2
ix0, iy0 = torch.maximum(px0, tx0), torch.maximum(py0, ty0)
ix1, iy1 = torch.minimum(px1, tx1), torch.minimum(py1, ty1)
inter = (ix1 - ix0).clamp(min=0) * (iy1 - iy0).clamp(min=0)
union = pw * ph + tw * th - inter
return inter / union.clamp(min=1e-8)
torch.manual_seed(0)
model = Detector()
opt = torch.optim.Adam(model.parameters(), lr=0.005)
Xt = torch.tensor(Xtr).unsqueeze(1); Bt = torch.tensor(Btr)
for _ in range(400):
opt.zero_grad()
loss = F.mse_loss(model(Xt), Bt)
loss.backward()
opt.step()
with torch.no_grad():
pred_te = model(torch.tensor(Xte).unsqueeze(1))
ious = iou_batch(pred_te, torch.tensor(Bte))
print(f'mean IoU on test set: {ious.mean().item():.3f}')
print(f'fraction of test boxes with IoU > 0.5: {(ious > 0.5).float().mean().item():.1%}')
fig, axes = plt.subplots(1, 4, figsize=(9, 2.5))
for i, ax in enumerate(axes):
ax.imshow(Xte[i], cmap='gray')
tb = Bte[i] * SIZE
pb = pred_te[i].numpy() * SIZE
ax.add_patch(patches.Rectangle((tb[0] - tb[2] / 2, tb[1] - tb[3] / 2), tb[2], tb[3], edgecolor='lime', facecolor='none', linewidth=2, label='true'))
ax.add_patch(patches.Rectangle((pb[0] - pb[2] / 2, pb[1] - pb[3] / 2), pb[2], pb[3], edgecolor='red', facecolor='none', linewidth=1.5, linestyle='--', label='pred'))
ax.set_title(f'IoU={ious[i]:.2f}', fontsize=9)
ax.axis('off')
axes[0].legend(fontsize=6, loc='upper left')
plt.show()
This detector above only handles exactly one object per image, because a fixed-size output vector (4 numbers) can only describe one box. But real scenes have a variable, unknown number of objects. Two different families of detectors dominate the landscape:
Two-stage (R-CNN family): first generate a modest number of region proposals — candidate boxes likely to contain something, via a cheap, class-agnostic method (the original R-CNN used classical segmentation; Faster R-CNN (Ren et al., 2015★) learns a small "region proposal network" instead) — then run a classifier-plus-box-regressor (this lesson's whole architecture) on each proposal independently, exactly like running the sliding-window classifier from Lesson 40 but only at a handful of promising locations instead of every window. Accurate, but only as fast as (proposals) x (one forward pass) allows.
Single-stage (YOLO, SSD): skip proposals entirely. YOLO (You Only Look Once, Redmon et al., 2016★) divides the image into a coarse grid of cells, and has each grid cell directly predict (as this lesson's network does) a fixed number of boxes plus a class label plus a confidence score, all in one forward pass. SSD (Single Shot MultiBox Detector, Liu et al., 2016★) follows the same one-pass recipe but predicts boxes from several feature-map resolutions at once (not just one final grid), so coarser layers naturally catch larger objects and finer layers catch smaller ones. To let a single cell describe objects of different aspect ratios, single-stage detectors use anchor boxes: several predefined box shapes (tall, wide, square) per cell, with the network predicting an offset from each anchor rather than a box from scratch. Faster to run, historically somewhat less accurate than two-stage methods, though the gap has narrowed considerably.
Both families end with the same postprocessing step: non-maximum suppression (NMS, Lesson 40) to merge the overlapping candidate boxes any real multi-object scene produces. (For simplicity, this lesson omitted NMS by only predicting a single bounding box.)
OpenCV ships a single-stage detector, cv2.FaceDetectorYN (YuNet, Wu et al., 2023), that follows the single-stage recipe above for faces specifically. Unlike Lesson 40's Viola-Jones cascade, it's a small single-shot CNN — the same family as SSD above, just specialized to one class — and it runs its own NMS internally before returning boxes.
import urllib.request
from pathlib import Path
import cv2
CACHE_DIR = Path.home() / '.cache' / 'cvintro'
YUNET_URL = 'https://media.githubusercontent.com/media/opencv/opencv_zoo/main/models/face_detection_yunet/face_detection_yunet_2023mar.onnx'
YUNET_PATH = CACHE_DIR / 'face_detection_yunet_2023mar.onnx'
def ensure_yunet():
if YUNET_PATH.exists():
return
CACHE_DIR.mkdir(parents=True, exist_ok=True)
print('Downloading the YuNet face detector (one-time, cached under ~/.cache/cvintro)...')
urllib.request.urlretrieve(YUNET_URL, YUNET_PATH)
ensure_yunet()
photo = cv2.imread('../img/apollo11_crew.jpg') # same photo Lesson 40 ran Viola-Jones on
h, w = photo.shape[:2]
yunet = cv2.FaceDetectorYN_create(str(YUNET_PATH), '', (w, h))
_, yunet_faces = yunet.detect(photo)
photo_rgb = cv2.cvtColor(photo, cv2.COLOR_BGR2RGB)
fig, ax = plt.subplots(figsize=(8, 6))
ax.imshow(photo_rgb)
for x, y, fw, fh, *_, score in yunet_faces:
ax.add_patch(patches.Rectangle((x, y), fw, fh, edgecolor='lime', facecolor='none', linewidth=2))
ax.text(x, y - 8, f'{score:.2f}', color='lime', fontsize=9, weight='bold')
ax.set_title(f'cv2.FaceDetectorYN: {len(yunet_faces)} detections')
ax.axis('off')
plt.show()
All three astronauts are found correctly, with no false positive — which is an improvement over Lesson 40's Viola-Jones cascade. This is the payoff of a learned single-shot CNN over a cascade of hand-designed Haar-like features: richer features, trained end-to-end on real face/non-face data rather than assembled stage by stage.
But YuNet only answers "face or not" — one class.
For general, multi-class detection, Faster R-CNN pretrained on COCO (Lin et al., 2014★ — about 330k real photos labeled across 80 object categories) is a few lines away via torchvision. This is the same "pretrained model in one line" pattern as Lesson 38's ResNet-18. It natively handles a variable, unknown number of objects per image — the actual payoff of the region-proposal architecture described above.
import urllib.request
from pathlib import Path
import cv2
import torchvision
CACHE_DIR = Path.home() / '.cache' / 'cvintro'
COCO_IMG_URL = 'http://images.cocodataset.org/val2017/000000039769.jpg'
COCO_IMG_PATH = CACHE_DIR / 'coco_sample.jpg'
def ensure_coco_sample():
if COCO_IMG_PATH.exists():
return
CACHE_DIR.mkdir(parents=True, exist_ok=True)
print('Downloading a COCO val2017 sample image (one-time, cached under ~/.cache/cvintro)...')
urllib.request.urlretrieve(COCO_IMG_URL, COCO_IMG_PATH)
ensure_coco_sample()
weights = torchvision.models.detection.FasterRCNN_ResNet50_FPN_Weights.COCO_V1
real_detector = torchvision.models.detection.fasterrcnn_resnet50_fpn(weights=weights)
real_detector.eval() # frozen, no training at all -- Faster R-CNN exactly as released
coco_classes = weights.meta['categories']
img_bgr = cv2.imread(str(COCO_IMG_PATH))
img_rgb = cv2.cvtColor(img_bgr, cv2.COLOR_BGR2RGB)
img_tensor = torch.tensor(img_rgb / 255.0, dtype=torch.float32).permute(2, 0, 1)
with torch.no_grad():
result = real_detector([img_tensor])[0]
score_thresh = 0.7
keep = result['scores'] > score_thresh
fig, ax = plt.subplots(figsize=(7, 5.5))
ax.imshow(img_rgb)
for box, label, score in zip(result['boxes'][keep], result['labels'][keep], result['scores'][keep]):
x0, y0, x1, y1 = box.numpy()
ax.add_patch(patches.Rectangle((x0, y0), x1 - x0, y1 - y0, edgecolor='lime', facecolor='none', linewidth=2))
ax.text(x0, y0 - 4, f'{coco_classes[label]} {score:.2f}', color='lime', fontsize=9, weight='bold')
ax.set_title(f'Faster R-CNN pretrained on COCO ({int(keep.sum())} detections above score {score_thresh})')
ax.axis('off')
plt.show()
Image source: COCO dataset (val2017, image 000000039769)
Four confident, correct detections survive the score_thresh=0.7 cutoff — two cats and two remote controls, each above 0.78 — despite this network using only pretrained weights. Lower score_thresh and rerun to see what the cutoff was hiding: several overlapping, lower-confidence boxes (a "couch" and a "bed" guess both covering most of the image, around score 0.54).
obj_size in make_scene from a fixed 8 to a random value (e.g. rng.integers(4, 12)) so objects vary in size, and retrain. Does mean IoU hold up, get worse, or barely change — and why would variable object scale be harder for a single fixed-size regression head than variable position?(cx, cy, w, h), but the metric that matters is IoU. Replace the loss with 1 - iou_batch(pred, target).mean() (directly optimizing IoU) and compare final mean test IoU to the MSE-trained version. Real detectors (e.g. Faster R-CNN, YOLO variants) do exactly this with generalized IoU losses — can you see why MSE loss and IoU metric might disagree on which of two similar predictions is "better"?score_thresh on the real Faster R-CNN to 0.3 and rerun. Count how many boxes appear versus at 0.7; which of the new ones are genuinely additional objects and which are duplicates/near-misses on the cats or remotes that NMS would normally merge away? Then try a different COCO val2017 image URL (swap the filename in COCO_IMG_URL, e.g. any http://images.cocodataset.org/val2017/<12-digit-id>.jpg) and see what classes the model finds.