Autonomous vehicles have the potential to reduce deaths and injuries dramatically and increasing the time to get there will cause thousands to millions of more deaths and an order of magnitude more injuries, not to mention property damage.
(I have no idea where lidar for self-driving cars falls within that balance, although I can't help but notice that human drivers do not utilize lidar.)
I suspect that right now, ML based approaches to depth perception would not pass the test to standards we would like, in practice requiring lidar. But, this would also provide an option to switch to an ML based gets good enough.
From a driver responsibility point of view, it is a fancy cruise control, there is no question about autonomous certification because the car isn't autonomous.
If one day, Tesla thinks it is OK for drivers not to pay attention and wants to make it legal, that's when the question about what kind of tech is reasonable will be asked. But for now, it looks like we are far from it. There is a huge gap between "works well most of the time" and full autonomy.
If a human has to sit there and monitor the car without interacting with it, then we shouldn't allow it on the roads. That means skipping level 2 and probably level 3. Go the Waymo route and go straight to full self driving cars.
Whether this is false advertising or even just moral is up in the air, but Tesla is making no attempts at educating consumers about the limits of their autonomy other than the bare minimum (yeah you technically have to keep both hands on the wheel, but that mechanism is easily defeated and you can even buy clips on Amazon that do it for you). I'm pretty sure Elon even liked a tweet about a Pornhub video where someone banged a model on the highway with autopilot. Which is an excellent day for PR and a disastrous day for legal.
The point being, sure, academically they are a far way away, but they are trying to convince consumers Level 5 is just around the corner.
Great for detecting if an object is in your way - but it cannot read a stop sign, you still need a camera system for that - so they still have to harden it against this hack (and who knows how many to come).
True, but LIDAR should make that detection more robust.
For example, if a stop sign is on the same plane as the object that surrounds it, that's a strong signal that it is not a real stop sign but rather a projected or displayed one.
That's much harder to detect with computer vision alone, especially if they're dealing with hackers who can test it with a real car and continually iterate on it to find different ways to trick it.
It's conceivable that in denser or older areas, signs are sometimes placed on structures because there isn't room for a pole.
Apple Maps shows stop signs and traffic lights on the map. A database like that would be another strong signal, but also not always accurate.
It makes me think that the first real wide scale self driving vehicle deployment will have to come from a nation state rather than a corporation; the easiest way to solve all of these problems is by putting heavy investments into smart road infrastructure such that the car depends on sensors embedded in the signs or the asphalt rather than any camera feed. A school bus could just have a transmitter to indicate whether its lights are on or not; construction workers could be given road signs that transmit similar signals or better yet use a portal to construct a virtual geofence around wherever the construction is happening before it begins. These are all easily solvable problems in today's world; what is missing is the political capital to invest in these solutions despite the insane ROI.
This infrastructure will need distributed two-phase commit. For safety-critical issues like school buses, eventual consensus won't be good enough. And then you'll need an infrastructure that detects when an older or simply defective car is present but not participating in the protocol. These problems are all solvable, but they're by no means trivial.
That is exactly why they can't interact with real world cases of fully valid plastic signs coming out of nowhere.
Still I wonder if it has failure modes that humans lack, which could result in situations where it does inexplicable things.
Almost certainly :)
This isn't a perception issue. It's an insufficient model for the world, and the software is obviously biased towards obeying critical traffic controls at lower confidence levels. It would be interesting to see what would happen if a (higher) speed limit sign was flashed in the same way.
Maybe .... before acting on a Stop sign sensed by the cameras the control system would confirm, via LIDAR, there is a post with a plate mounted on it in a position which corresponds to where the camera thought it saw a Stop sign ?
Can LIDAR do that level of sensing, ie unambiguously find a metal post/plate out of all the other stuff which might be in front of the car ?
And assuming my guess about how LIDAR would help is correct what happens when an obstruction between the car and the Stop sign means that LIDAR can't 'see' the post but the cameras can see the sign ?
You can of course pick out discrete examples of humans failing tasks like these, but entire classes of computers routinely fail at things like this.
Well surely that's what we should be testing. Is a self-driving car system safer than a human, or not? You're correct that self-driving car systems lack a "general learning and intelligence model," but they also have several obvious advantages over humans, like potentially better-positioned cameras, a more reliable "attention" system, faster reaction times, potentially more "experience" with different road conditions and scenarios, etc.
Furthermore, everything about the road is designed to help us humans. The size, color, placement, etc... of road signs and markings is specially designed for human brains to process.
As for the hardware, our eyes have very good resolution, field of view and an unmatched dynamic range. Far better than the stuff in Teslas.
So I'm not saying that it is impossible for computer vision to get away with just using cameras, but it is competing with what our brain does best. So a little help in a form of a LIDAR isn't too much IMHO.
Humans do get very clunky and slow when that depth perception is taken away. One example is night vision goggles which basically just project a flat screen onto your eyeballs. Many people report that they become susceptible to tripping, falling, hitting objects, or just straight up walking into trees with NVGs, in part because the loss of depth perception takes such a toll on their brain that they can't adapt in time. It takes hours if not days to fully adapt to NVGs, and even then you can't operate as quickly as you can in daylight unless you practice consistently (think SWAT or Navy Seals type training).
Cameras lack any perception; perception is a function of brains.
But a single camera with appropriate processing can provide a wide array of depth cues (stereoscopy adds more.)
> Many people report that they become susceptible to tripping, falling, hitting objects, or just straight up walking into trees with NVGs, in part because the loss of depth perception takes such a toll on their brain that they can't adapt in time.
While loss of depth cues plays some role here, a much bigger factor, AIUI, is the narrow field of view. “Panoramic” NVGs have around 100° horizontal and 40° vertical field of view (non-panoramic NVGs often have around a 40° horizontal FoV), people have around (with eye but not head movement) a 210° horizontal and 120° vertical field of view. Looking directly ahead with even panoramic NVGs as you normally would without them, you literally cannot see lots of things, especially tripping hazards, and objects to your sides that you might laterally move into to avoid the things you can see in front of you.
The reason people become clumsy when using night vision devices is because they have a tiny field of view that only moves when your head moves. You can close one eye and walk around just fine. But if you attach two toilet paper tubes to your face, you'll have trouble balancing and avoiding obstacles. The clumsiness is caused by the lack of peripheral vision and the inability to quickly glance (since the only way to change what you see is to turn your head).