. How does this lesson's detector represent the box it predicts?
The network's regression head outputs (cx, cy, w, h) through a sigmoid, keeping every value in the valid normalized [0, 1] range.
. Why is IoU used to evaluate the detector instead of just looking at how close the four predicted numbers are to the true ones?
Two boxes can have similar-looking coordinates but very different overlap (or vice versa), so IoU is the metric that reflects what actually matters for localization.
. Why can this lesson's simple regression detector only ever predict exactly one object per image?
A fixed 4-number output has no way to represent a variable number of objects; real multi-object detectors need a different architecture (grid cells, region proposals, etc.).
. What is the key architectural difference between two-stage (R-CNN family) and single-stage (YOLO, SSD) detectors?
Two-stage methods (e.g. Faster R-CNN) propose-then-classify; single-stage methods (YOLO, SSD) predict boxes and classes directly per grid cell in a single pass, trading some accuracy for speed.