← Back to Lesson 33

Quiz 33: Optimization

Score: 0 / 4

. On the narrow, steep-in-y bowl loss surface, why does plain gradient descent zigzag instead of converging smoothly?

A single shared step size can't be simultaneously right for a shallow and a steep direction at once, causing the characteristic zigzag.

. What is the key mechanism behind momentum's benefit on the narrow bowl?

Momentum's real benefit is acceleration along consistently-pointing directions, not damping oscillation, since a single beta can't be tuned per-axis.

. How does Adam differ from momentum in how it handles the steep vs. shallow axes of the bowl?

Dividing by sqrt(v) per-parameter automatically shrinks steps on the steep axis and boosts them on the shallow one, unlike momentum's single shared coefficient.

. Why does initializing all weights to exactly zero fail catastrophically for a hidden layer with more than one unit?

With identical weights and identical gradients, every unit updates identically forever, so the model is stuck computing a single trivial function no matter how long it trains.