I hear this comment a lot when defending Tesla's choices, and it's a red herring. The fact that humans only rely on two cameras means nothing. Repeating old comments of mine:
You also don't "need" megawatts of power to play top-level Go: humans do it with 100 watts of energy. Yet Google needed who knows how many megwatts of energy to train and run AlphaGo on their massive server farm.
Imagine two companies competing to win at Go, and one company had the attitude that megawatts of energy was not necessary for training and prediction, and another company threw the biggest GPU farm they could. The second company just played top-level Go this year. The first company is ~10 years away from a low-energy elite Go computer.
Humans implicitly perform SLAM (simulataneous localization and mapping). What do I mean? Look around your room. Close your eyes. Visualize the room. As a human, you've built a rough 3D model of the room. And if you keep your eyes open and walk through the room, that map is pretty fine-grained/detailed too and humans can keep track of where they are in the map.
Doing this accurately in moving environments (especially with lots of pure forward motion) with just two cameras is still a wide open research problem.
Doing this with LIDAR/GPS/IMU/recorded maps is solved. That's why people use LIDAR.
Matching the abilities of human perception is an insanely hard problem. Don't let cute problems like image classification fool you. Why make the problem even harder?