← Back to Lesson 31

Quiz 31: Neural Network Fundamentals

Score: 0 / 4

. Why does the lesson use the sigmoid function instead of directly thresholding the raw score z = w^T x + b for training?

Hard accuracy jumps discontinuously at the boundary with zero gradient almost everywhere, so gradient descent has nothing to follow; sigmoid squashes the score into a smooth probability instead.

. In binary cross-entropy, what happens to the loss when the true label is 1 but the predicted probability p is close to 0?

The penalty for true label 1 is -log(p), which grows without bound as p approaches 0.

. What two independent methods does the lesson use to verify the hand-derived backprop gradient formulas are correct?

The lesson perturbs each parameter by a tiny amount to estimate the gradient numerically, and separately lets PyTorch's .backward() compute it automatically -- both agree with the hand-derived formula.

. Why does gradient descent do *worse* than exhaustive search on the ring-vs-disk dataset, converging to near-chance accuracy?

Because the ring and disk are both centered at the origin, the training signal from all points nearly cancels in every direction, so the optimizer settles near w=0 instead of finding an asymmetric corner-case solution.