that is good experimental data and the role of theorists here is to look how that performance achieved and why. For example, one can reasonably suspect that there is a good reason why the kernels in a well trained image recognition deep learning net do look like receptive fields of neurons in visual cortex. I'm pretty sure that there is some kind of statistical optimality in that, something similar to like normal distribution is maximum entropy distribution for a given standard variation. The same way i'd guess Gabor of neuron receptive field is something like maximum entropy on the set of all possible edges or something like this. The point here is that the great success of deep learning generates a lot of very good data for theorists to consume. You can do only so much theory without good experimental data, and in the decades before the availability of computing power (and resulting success of deep learning) there wasn't that much of the computer vision theory advances to speak about, really.
>leaves flying and various things
Newton did that for 20 years. With great success.