Elon's objections to lidar are not because it's a bad idea but because of the costs. Especially because he's already sold cars with "full self driving" that don't have LiDAR. So he's painted himself in a corner.
Even if it can be done with vision only, lidar will add an extra data point for cases where the object recognition is in doubt. And the price will come down.
One should think of the car cameras as a fixed periscope view. Try driving with that, what the computer sees is more like driving a tank with monitors. Not the best.
It’s only reasonable to give the cars some super human senses to compensate, like LIDAR etc.
They held out longer than expected, some 20 years or so, but common sense prevailed. Even if they had to make a design where the buttons weren't clearly visible first.
Even today Macs don't really have 2 mouse buttons, they just call a double-finger click a "gesture" even though it is simply a right-click.
Part of the reason I've owned FSD since 2017 and yet spent exactly zero minutes in my car in the expressway, in light traffic, in good weather, in the middle of the day reading a book is because of the lack of an entry on the posted timeline saying something like: "Tesla states it will accept liability for any crashes that occur while FSD was engaged."
After all these years, it's fair to say: "In the 7 years that Tesla has accepted money from customers for FSD, Tesla still requires human supervision in all cases."
So far they have proven they are asmpytotically approaching something that drives badly and is in danger of hitting things.
I think the Pure vision stuff is nifty, but “nifty” is the exact dead-last parameter when trying to transport my family safely from point A to point B.
Oh, it can't?
So you need cameras too?
So you have to build a 3-dimensional sensor fusion system for LIDAR to work?
Wouldn't that fusion system be more complex, less performant, and more fallible that just choosing a single model (vision/LIDAR) and optimising around that?
10 million years of animal evolution makes a good case for stereoscopic vision as a sensing system for navigating the world.
LIDAR is a crutch for low-maturity software systems.
Stereoscopic vision? That's when you apply sensor fusion to two cameras (and a number of microphones, too), isn't it?
If airplane designers had trusted in "animal evolution", planes would flap their wings to gain height.
Overall, the complexity of modeling the world the way an animal does seems much much bigger than a few different sensors.
Adding "a few different sensors" is not going to solve the fundamental problems of navigating the world.
If Tesla had been limiting its ideas to theoretical research, or even applied research, I'd be all for vision only as a valid research avenue. But they are putting this thing on the streets where I walk, and they are charging people thousands of dollars with claims that it works today, and that it will do wonders tomorrow. That is simply not acceptable for a green field research concept that probably has decades left in front of it.
No, the whole point about sensor fusion is that it can be greater than the sum of it's parts.
My whole point is that a single system is not sufficiently safe to do what we're trying to do. The point of FSD isn't just to navigate without hitting stuff. The point is to do it with ~99.999% accuracy, ~99.999% of the time while flying down the road at 80 MPH.
Humans are _terrible_ at this, just check the deaths due to automobile accidents each year. We have some pretty amazing stereoscopic vision. But I don't like trusting my safety to another human, and I sure won't trust it to a machine whose vision isn't as good as mine.
For me to trust a machine, it needs to be an order of magnitude better than myself.
It is actually shocking that anybody pushes the "good enough for evolution" narrative when Tesla has completely and utterly failed to do even that right. This is even ignoring all of the other purely mechanical inadequacies of cameras relative to eyes such as resolution, dynamic range, dynamic attention-based focal lengths, being mounted on mobile swivel to allow for parallax calculations, etc. Let alone the other neurological elements that are not fully understood of integral to human-level perception. But no, they could not even get the two eyes instead of one eye part right.
but yeah, broadside of a semi mistaken for the sky
Your best chance today is to pack more sensores.
Edit: I still believe that the systems which assist the driver today are useful and can make driving safer and easier. I don't want to downplay what they have achieved but they are a looong way from L5.
The problem isn't just training, it's sensors. There are simply no sensors available today that work in all the different weather conditions that L5 requires with the required precision.
Tech is not biology, so in order for a computer to not do something stupid, it generally needs different sensors. Some sensors (depth gauge when scuba diving, altimeters for planes, or even gps) are obvious examples of when tech outperforms humans and is safety critical. Just because humans can drive doesn’t mean computers should replicate our “control loop”.
Our distance estimates are quite good at human scales, especially when compared to vision based approaches. We don't have great precision (we will never say something like "it's 19.7m away") but we very rarely grossly misestimate (think something is 10m away when it's actually 200m away or 10cm away).
Of course, compared to a LIDAR or advanced techniques used in physics, we don't hold the tiniest candles (say, if we were to compare to LIGO's distance measurements).
Not using a "crazy number of sensors" is not something you can brag about outside a kind of "sell me this pen" sales pitch.
This can be true without both ever materializing though ;)