. In the overfitting demonstration with only 12 training images, what pattern do the training and validation loss curves show?
The network perfectly memorizes the 12 training images while validation loss reaches a minimum then gets worse as the model overfits further.
. How does data augmentation help fight overfitting on a small dataset?
Applying random label-preserving transforms to the existing images gives the model many more effective examples to learn from, rather than memorizing a handful of exact images.
. What does weight decay do to fight overfitting?
Penalizing large weight magnitudes makes memorizing noisy specifics of a small dataset more costly relative to finding a simpler, smoother function.
. In inverted dropout, what happens to the surviving (non-zeroed) activations during training, and why?
Rescaling by 1/(1-p) keeps the expected activation magnitude consistent between training (with dropout) and evaluation (without it).