Score: 0 / 4
. Why does the lesson use the sigmoid function instead of directly thresholding the raw score z = w^T x + b for training?
Hard accuracy jumps discontinuously at the boundary with zero gradient almost everywhere, so gradient descent has nothing to follow; sigmoid squashes the score into a smooth probability instead.
. In binary cross-entropy, what happens to the loss when the true label is 1 but the predicted probability p is close to 0?
The penalty for true label 1 is -log(p), which grows without bound as p approaches 0.
. What two independent methods does the lesson use to verify the hand-derived backprop gradient formulas are correct?
The lesson perturbs each parameter by a tiny amount to estimate the gradient numerically, and separately lets PyTorch's .backward() compute it automatically -- both agree with the hand-derived formula.
. Why does gradient descent do *worse* than exhaustive search on the ring-vs-disk dataset, converging to near-chance accuracy?
Because the ring and disk are both centered at the origin, the training signal from all points nearly cancels in every direction, so the optimizer settles near w=0 instead of finding an asymmetric corner-case solution.