. Why does stacking two purely linear layers (no nonlinearity in between) fail to add any representational power over one layer?
Without a nonlinearity, z = w2^T(W1 x + b1) + b2 simplifies algebraically into a single w'^T x + b', exactly Lesson 31's single neuron again.
. What role does the ReLU nonlinearity play in an MLP's hidden layer?
Because ReLU is not a linear function, w2^T ReLU(W1 x + b1) + b2 cannot be rewritten as any single linear projection, unlike the all-linear case.
. In the capacity sweep, what does H=1 (a single hidden unit) achieve on the ring-vs-disk problem, no matter how long it trains?
One hidden unit provides only a single straight-line cut, so it inherits the same representational ceiling a lone linear projection has -- a capacity limit, not an optimization failure.
. When the trained hidden layer's 4D output is projected to 2D with PCA for visualization, what does the lesson observe about the two classes?
The hidden layer reshapes the space so the classes become nearly linearly separable, illustrating that each layer's job is to make the next layer's job easier.