. On the narrow, steep-in-y bowl loss surface, why does plain gradient descent zigzag instead of converging smoothly?
A single shared step size can't be simultaneously right for a shallow and a steep direction at once, causing the characteristic zigzag.
. What is the key mechanism behind momentum's benefit on the narrow bowl?
Momentum's real benefit is acceleration along consistently-pointing directions, not damping oscillation, since a single beta can't be tuned per-axis.
. How does Adam differ from momentum in how it handles the steep vs. shallow axes of the bowl?
Dividing by sqrt(v) per-parameter automatically shrinks steps on the steep axis and boosts them on the shallow one, unlike momentum's single shared coefficient.
. Why does initializing all weights to exactly zero fail catastrophically for a hidden layer with more than one unit?
With identical weights and identical gradients, every unit updates identically forever, so the model is stuck computing a single trivial function no matter how long it trains.