That "some kind of image processing unit" in humans has an awful lot of compute power and software.
If you remove $100k of sensors but have to add $200k of compute to run more advanced computer vision software, then it's a bad tradeoff to use only cameras, even if in theory that software is possible.
And yeah, as you mention, cameras don't really have the same level of range our eyes have and computers don't operate in the same way.
I think more practically cars have adding driver assistance feature for a while now - more cameras, blind spot monitoring, ultrasound for parking, lane drift indicators.
It is therefore not unreasonable to assume that adding more sensors is helpful (but even the old adage of more data is better than less would probably say that).
I'm not a fan of the camera-only approach and think Tesla is making a mistake backing it due to path-dependence, but when we're _only_ talking about this is _broadly theoretical_ terms, I don't think they're wrong. The ideal autonomous driving agent is like a perfect monday morning quarterback who gets to look at every failure and say "see, what you should have done here was..." and it seems like it might well both have enough information and be able too see enough cases to meet some desirable standard of safety. In theory. In practice, maybe they just can't get enough accuracy or something.
In certain conditions, yes. Humans drive terribly in dark and low light, something lidar excels in.