. What is the fundamental difference between semantic segmentation and detection (Lesson 41)?
Detection localizes objects with boxes; semantic segmentation goes further and assigns a class label to every single pixel in the image.
. In this lesson's minimal FCN-style network, what does the decoder see when producing its output?
A plain FCN just downsamples then upsamples: the decoder has no access to the encoder's higher-resolution intermediate features, only the coarse bottleneck.
. How do U-Net's skip connections fix the boundary-detail loss of a plain FCN?
Skip connections concatenate encoder features into the decoder at each matching resolution, so fine spatial detail never has to survive the lossy bottleneck in the first place.
. When the real pretrained FCN is run on a photo containing a tram, it labels part of the tram `train` and part `bus`. What does this reveal?
The split roughly tracks a real visual seam (windows vs. body) and reflects VOC's limited vocabulary, not a random or nonsensical failure — a reminder that a model's answers are only as good as its label set.