This shouldn't really be surprising. Machine learning is specifically not magic. The reason CNNs have seen so much success is precisely because they build in translation-invariance, which massively cuts down on parameters while forcing the final function to have the desired structure regardless of wherever gradient descent takes the weights.
Also why most papers in deep learning are network architecture innovation.