A pixel with large gradient magnitude is one the network's prediction is most sensitive to -- a pixel it's 'looking at'.
. Why are Grad-CAM heatmaps coarser (lower-resolution) than saliency maps?
Grad-CAM's resolution matches whichever conv layer it uses, which is coarser than the full input resolution a saliency map operates at.
. On *clean* test images (no marker present), why does Grad-CAM still catch model_shortcut's reliance on the marker far more reliably than raw saliency does?
The lesson reports Grad-CAM's peak lands in the marker region 92% of the time on clean images versus only 56% for saliency, attributed to Grad-CAM's channel-level pooling smoothing out pixel noise.
. What does t-SNE do differently from a linear projection like PCA (Lesson 6)?
t-SNE emphasizes preserving which points are close to which other points locally, which is why same-class images can end up visibly clustered even without ever seeing their labels.