There is also an information theoretic lens about compression, which is probably very close to what you are thinking about, but I haven't studied it yet.
Regarding Image -> Animal. An image of an animal is a projection of the animal onto a 2D plane, plus lots of noise. So there is some dependance between an image of an animal and the animal. Biologists can get a lot of information looking at a photo of an animal. In some sense they are always looking at two images from their eyes.
But the problem you are talking about is indeed serious, and far from solved. You can't understand the real world from 2D images with the current approaches. Ideally we want neural networks to build a 3D (or even 4D, with time?) model of reality. Instead we find them trying to guess labels based on patterns. May favourite example is the tiger-dog [1]. Still, there is evidence NNs are doing some clever things [2]. My guess is that the problem is that we just haven't found a way to formulate the task for the solution we want. In the current formulation it's easiest for the model to minimize the loss by sticking to patterns, so why do something else?
There is a lot of research on more applied ML that asks the questions you are asking. It's just that this paper is another attempt at a theoretical explanation.
I agree on the self-driving cars. We can't have truly self-driving cars until a model can generalize, which none can't at the moment. The core question is: if we had a "dumb" model that does clever averaging and successfully covers 99% cases, such that the car is dumb in special cases, but smarter than most human drivers in usual cases, would it justify deploying the cars? If this was the case, dumb ML might be enough for self-driving. It's definitely enough for self-driving in walled garden conditions, so there is some evidence that with enough data we can brute force our way to a tolerable solution.
[1] https://www.dropbox.com/s/ucvflwwrm8idnp6/photo_2022-03-29_2... [2] https://distill.pub/2020/circuits/