Currently CV can be good or fast, but not really both.
Also, contrariwise, your brain is actually terrible at depth perception and object recognition, and employs many tricks to approximate depth perception and recognition[0]
---
I don't have any special knowledge on this, but I wouldn't be surprised if self-driving car companies are using Lidar to train CV algorithms.
---
0. http://www.opticalillusionsportal.com/wp-content/uploads/201...
The human eye does the equivalent of what you might call supersampling in computer terms. The eyes make very subtle focus shifts to change perception of the same object and get a more detailed picture of something that is difficult to see. No mechanical lens that I know of can move this fast or accurately but you can achieve something similar with multiple cameras.
The brain also filters out extraneous data and interpolates missing data.
Human vision is far from perfect though. Light changes that exceed you sensitivity range will blind you. Vision is impaired by darkness.
The filtering and interpolation mechanisms of the brain are error prone and the brain often does not know when it is wrong. You may see something and never know it wasn't real, or don't see something and never know it was there.
But it's hard to say necessarily that it's not optimal -- that is, that a digital version could necessarily do better. We are biased by our priors, just as a digital pattern matcher is biased by its training set. And in low light, we try to find patterns that we recognize in the noise, reconstruct the missing parts, etc., just as a computer would.
A digital system might eventually have better resolution and be able to do better on some scale of precision and performance, but I suspect that most attempts to improve performance by introducing "supersampling" and pattern matching and other forms of inference will always result in similar errors to the brain. Perhaps different in character due to differences in the training set and algorithms, but of a similar nature.
And those human brains are attached to human arms which navigate Americans into 32,000 motor vehicle deaths a year [1].
[1] https://en.wikipedia.org/wiki/List_of_motor_vehicle_deaths_i...
Or our ability to interpret facial expressions of pedestrians to help gauge if they might walk out in front of the car.
Why make the problem harder than it needs to be? Why not shoot for superhuman? Who cares if the submarine swims like a fish?
We add ultrasound range sensors for parking because it helps, even though it's possible to park despite not being bats. You also could still drive with reduced vision, but it'd be better if you didn't.
Would we drive better if we had accurate distance sensors?
So we theoretically may not need to, but they help with the current problems we have.
Problem with ultrasonics is latency, range as well as lack of accuracy.
To detect obstacles with vision, generally you need stereo camera and generate disparity map (you can do with moving mono camera too but its much harder and even more unreliable). The confidence levels drops quickly as objects are farther and lighting conditions, reflections etc gets more complex. The main problem is that popular algorithms only rely on geometry to compute the depth and they don't into account knowledge of the world. For example, if there is area of specular reflection, geometry will fail even though one can reasonably estimate depth by assuming some continuity of shape of objects in real world.
We humans are very efficient in figuring out depth even with single eye and its purely because we can do very quick and accurate object segmentation and combine it with our knowledge of what those objects are (semantics) as well as our internal map of the world around us. For example, when you see truck on the road even with single eye, you have fair estimate of what sizes trucks usually are and how big should it appear if it was 100ft away vs 1000ft away. You would be able to do this even if you can see truck only partially from the side angle or back angle.
There is no known algorithms currently that can come anywhere close to human level accuracy and speed for depth estimation and consequently obstacle detection using just vision in variety of lighting conditions when you are moving at 70 miles/hr. On the other hand, lidars are super easy to implement, fast and very accurate. So unless a major breakthrough comes along, lidars are actually the only choice right now for fast and accurate obstacle detection. Remember even a tiny error rate could be fatal when you have millions of cars covering millions of miles every day.