Train it with as many images as you want and as long as a good enough face shows up, the model is going to have a positive match. The entire problem is it’s missing that upper level of intelligence that evaluates “that looks like a face, could it actually be a human?”
Is there? Humans used to think the gods were literally watching them from the sky and the constellations were actual creatures sent into the night. So this seems learned behavior from data rather than some inherent part of human thinking.
>Train it with as many images as you want and as long as a good enough face shows up, the model is going to have a positive match.
So will a human if something is close enough to a face. A shadow at night for example might look just like a human face. Children will often think there's a monster in the room or under their bed.
Humans do not need to be trained on billions of images from around the globe to semantically understand where human faces are not expected to appear. Modern AI can certainly recognize faces really well with that level of training now, but it still doesn’t even understand what a face is (i.e. no model of reality to verify its identification against).
(I was thinking of this when I was driving in a new place. Suddenly it looked like the road ended abruptly and I got ready to act, but of course it didn't end and I realized that just a split second later.)
Pareidolia: https://en.wikipedia.org/wiki/Pareidolia
This really is a god-of-the-gaps answer to the concerns being raised.
For consideration, our brains start with architecture and connections that have evolved over a billion years (give or take) of training. Then we are exposed to a lifetime of embodied experience coming in through 5 (give or take) senses.
ML is picking out different things, but it's not obvious to me that models are actually getting more data then we have been trained on. Certainly GPT has seen more text, but I don't think that comparing that to a person's training is any more meaningful than saying we'll each encounter tens of thousands of hours of HD video during our training.