Interesting! Made me wonder - how would we compare the data that neural net receives with the data that toddler's visual cortex is getting?
On a purely visual level - let's say we have 10k of static 32x32 images defining a class of cat. Or even more of them plus some negative examples. Each image is a different cat, in a different position (they're incredibly flexible creatures). Having so many cases we should be trying to make some kind of a generalization of what 1024 pixels of a cat should look like.
A family with a toddler has only one cat. From that one example, he learns the concept of what a cat is and is able to generalize it when in different situations. But a toddler has 2 eyes, his visual input is stereo. Even if he sees just one cat, it's not a static image interaction. The input is temporal, he can see the cat moving and interacting with the environment.