Looking through a small window at a moving edge, only the motion perpendicular to the edge is visible; motion along the edge is invisible locally.
. In Lucas-Kanade, why does the flow estimate at a corner tend to be reliable while at an edge it is not?
Solving the Lucas-Kanade system requires M to be invertible; a corner (two large eigenvalues) is well-conditioned, while an edge (one small eigenvalue) is ill-conditioned -- the aperture problem again.
. Why does cv2.calcOpticalFlowPyrLK use an image pyramid and iterate, rather than solving once?
Estimating coarsely on a small, blurry pyramid level first, then refining level by level, lets large motions be recovered even though the linear approximation only holds locally, one small step at a time.
. What does the Horn-Schunck smoothness term let flow estimates do that Lucas-Kanade's purely local windows cannot?
The smoothness term couples every pixel's flow to its neighbors', letting information spread from regions with strong gradients into flat regions that have no data term of their own.