Pass: An ImageNet replacement for self-supervised pretraining without humans
robots.ox.ac.uk
robots.ox.ac.uk
One could easily argue that throughout all of human evolution, recognizing and interacting with other humans was the single most important factor for survival. Accordingly, one would expect the human visual system to be highly adapted towards recognizing humans. [1]
If we want AI to safely interact with humans, then we should strive to make the AI understandable in the sense that our intuition of what it might see or do matches up with what it'll actually see or do. That's a big aspect why those Tesla crashes are so scary: No human would overlook a huge white truck sideways on the road. But an AI trained on the wrong dataset might.
Accordingly, I would have expected that datasets move more towards imitating what babies will see in their first days or weeks, because that seems to be the most promising path towards replicating human vision with AI.
[1] https://news.stanford.edu/news/2012/december/infants-process...
This seems totally appropriate for pretraining on natural images IMO.
If an FSD vehicle is better at driving than the average driver, then I am safer, regardless of my driving ability, because it means other cars are less likely to crash into me.
If I am above average driver, I can make the choice to continue driving instead of letting FSD take the wheel, but I probably won't.