. What made the ILSVRC benchmark (built on a subset of ImageNet) so important for comparing architectures?
A shared, standardized competition -- not just a big labeled dataset -- is what let different architectures be compared fairly and drove rapid progress.
. According to the lesson, why does stacking more layers make a *plain* (non-residual) deep network worse, not better?
This is the vanishing-gradient problem: repeated multiplication by small per-layer factors drives the gradient toward zero, so early layers stop learning.
. Structurally, why does a residual connection (x_{l+1} = x_l + F(x_l)) prevent the gradient from vanishing across many layers?
The identity matrix in the Jacobian is never shrunk by a small activation derivative, unlike a plain network where every layer's Jacobian is purely dF/dx_l.
. What do the two learnable parameters gamma and beta do in batch normalization, after a channel's activations are normalized to mean 0 and standard deviation 1?
gamma (scale) and beta (shift) are applied on top of the 0/1-normalized activations, so the network can still recover a different mean/scale if that's actually useful.