Contrast helps. But running a convolution model on a single video frame, or aggregate of frames, is not how human vision works either. Give a human a glob of pixels with no context and they will struggle to identify it too. Identification involves motion and knowledge of the scene.