A couple of things:
* We have seen more than 100,000 "photos" in the sense that photos are just images - if photos are just images, we have a constant feed of "photos" every single moment our eyes are open. Of course, that's not the same as these training datasets, but it is still worth keeping in mind.
* All of these things trained on massive datasets with self-supervised learning are in a sense addressing the "slowness" of learning you mention, since self-supervised (aka no annotations are needed beyond the data itself) "pre-training" on the massive datasets can then enable training for downstream tasks with way less data.
* Arguably requiring massive datasets for pre-training is still a bit lame, but then again the 4-5 years of life it takes to reach pretty advanced intelligence in humans represents a whoooole lot of data. And as with self-supervised learning on these massive models, a lot of intelligence seems to come down to learning to predict the future from sensory input.
* Humans also come with a lot of pre-wiring done by evolution, whereas these models are trained from scratch. Evolutionary wiring represents its own sort of "pre-training", of course.
So basically, it is not so slow to learn as it seems. Arguably it could get faster once we train multimodal models and concepts from text can reinforce learning to understand images and so on, and people are working on it (eg GATO). There may also need to be a separation between low level 'instinct' intelligence and high-level 'reasoning' intelligence; AI still sucks at the second one.