> There are surprisingly sophisticated (and precise) machine learning algorithms for mono-camera depth estimation
There are, and I wouldn't trust them with life critical stuff. The more accurate ones are not realtime. The hardpart is that they are very noisy. unknown objects wobble about in depth considerably. You need to do lots of filtering to get useful results, which eats into time.
It is very much an unsolved problem. Its one that I'm partly working on now. However for the device that I"m partly working with, monocular estimation is far too power hungry, noisy and generally shite.
> Furthermore, most companies use various focal length lenses.
no, tesla _rely_ on having a wide, medium and zoom camera. However they blended into the same effective sensor to give a better chance at tracking objects. I bet you a $10 that if you block one, the whole system turns to shit.
> The real problem lies in sensor fusion.
Sensor fusion is trivial. Its understanding what they sensors are telling you, thats the hard part.
for example, when trying to turn across oncoming traffic, all the sensors will tell you that stuff is coming towards you, what they won't tell you is if its safe to cross. Thats the hard part. Given that tesla can't accurately place a car on the road yet, they can't safely cross traffic.
Sensor understanding is the problem that is really unsolved. Depth estimation with stereo cameras is 85% of the way there, monocular estimation is no where near that level, simply because robust object recognition isn't going to be a thing for at least another 6 years.