Or am I missing something?
Beating the prime of the prime human driver with cameras may be hard. Beating tired / drunk / old drivers is a different story.
I have an adversarial example.
Systems wise, that's going to be a lot more complex than having more cameras. Of course, if you've got enough cameras you can probably get better visibility than a human, and continuously process all viewpoints. But 2 cameras is definitely insufficient.
A camera or even a stereo pair of cameras mounted in or on the car will provide inferior imagery to the control system than eyes to the brain. They have less dynamic range and no articulation. If you wanted to replicate human style vision you'd need a bunch of fixed cameras and inertial and acceleration sensors all on top of AI that's better than what Tesla's been demonstrating.
LIDAR is the most straightforward augmentation for fixed cameras because it can build very accurate depth maps and image segmentation. You need fewer fixed cameras if your spacial model is built with LIDAR. You're in even better shape if those systems are augmented with radar.
While humans don't have LIDAR and such, our visual systems are highly developed and augmented with highly developed proprioception. Trying to replicate it with just cameras and tons of processing power is a fool's errand.
I remember the first time we had problems with a matte black surface with our LIDAR that would have been easily spotted by our camera and vice versa with a shiny white surface in direct sunlight relative the car being easily picked up by lidar but nearly invisible to the cameras.
In theory a camera-only autonomous system can work effectively but in practice there's innumerable edge cases where it doesn't work well and edge cases are where catastrophic failures happen. If you had infinite processing power and error free AI you might be able to cover many edge cases but Teslas have neither.
Now, there’s this little thing called The Pareto Principle.
That last 20% is going to take a long time to achieve, and is going to be very, very expensive.
Are you willing to roll a ten sided die every time you get in the car, and only if you roll a three or higher, do you get to arrive at your destination unhurt, on time, and without major incident?
No surprise then that the vision-only "level 5 autonomous driving" system would want the high beams on...
Stereo vision is mostly usable below 15 ft/5m range, if even that.
Stereo vision can also be fooled by things like reflections.
Stereo vision at the human level also requires head movements and so on.
And what about how fish use fins to swim?