← Back to Lesson 37

Quiz 37: Classic Architectures

Score: 0 / 4

. What made the ILSVRC benchmark (built on a subset of ImageNet) so important for comparing architectures?

A shared, standardized competition -- not just a big labeled dataset -- is what let different architectures be compared fairly and drove rapid progress.

. According to the lesson, why does stacking more layers make a *plain* (non-residual) deep network worse, not better?

This is the vanishing-gradient problem: repeated multiplication by small per-layer factors drives the gradient toward zero, so early layers stop learning.

. Structurally, why does a residual connection (x_{l+1} = x_l + F(x_l)) prevent the gradient from vanishing across many layers?

The identity matrix in the Jacobian is never shrunk by a small activation derivative, unlike a plain network where every layer's Jacobian is purely dF/dx_l.

. What do the two learnable parameters gamma and beta do in batch normalization, after a channel's activations are normalized to mean 0 and standard deviation 1?

gamma (scale) and beta (shift) are applied on top of the 0/1-normalized activations, so the network can still recover a different mean/scale if that's actually useful.