One more relevant note - (Olshausen and Field, 1997) showed that the filter employed by V1 simple cells could be learned using some simple assumptions about sparse coding and a single image. Translation invariance built in by way of the sampling scheme of the image, small patches.
The filters learned by the first layer of CNNs is usually of the same type, Gabor filters. Not a coincidence.
That was twenty years ago. What's old is new?