The author mentions correlation, but doesn't properly give the intuition of what's happening:
> When we perform convolution of an image of a person with an upside image of a face, then the result will be an image with one or multiple bright pixels at the location where the face was matched with the person.
Well yea, but the best way to explain it is that when you flip the signal and convolve, you're getting the largest summation when the two align. Simple as that.
I'm sure we'll see the Circular Harmonic Transform for rotation invariant (up to a point) Deep Nets, and maybe even the Mellon Transform for scale invariant (again, up to a point) Deep Nets. <- Note i haven't done the math per say showing that these will work, but i can't think of a reason they wouldn't.