. What two problems does a convolutional layer's weight-sharing fix, compared to a fully-connected layer applied to a flattened image?
Restricting each output to a local patch and sharing that patch's weights across positions both shrinks the parameter count and makes detections shift consistently with the input.
. For a 3x3 convolution filter applied to a 3-channel RGB image, how many weights does that single filter actually have?
A conv filter is really 3x3xC_in, so on 3 input channels a '3x3 filter' actually has 3x3x3=27 weights, summed into one output value per position.
. What does max pooling add to a CNN, beyond shrinking the spatial size of the feature map?
Keeping only the max value in each window means small position shifts within that window don't change the pooled output.
. In the unseen-position generalization test, why does the flatten-based MLP perform near chance on shapes placed at corner positions never seen in training, while the CNN does not?
Flattening ties the MLP's decision to specific pixel positions, while the CNN's global pooling collapses spatial position, letting it recognize shapes regardless of where they appear.