So it's just as much about how humans can reason from sight as it is about how good their sight actually is. It's very possible that by the time you've built a computer capable of the same sort of reasoning that humans do, you've actually solved, well, artifical general intelligence and you're no longer designing a car at all - because the AGI took that job ages ago.
Though I think that robotic automobiles must perform better than humans, so equipping them with better sensors makes perfect sense, regardless of algorithms.
Also I think that it actually proves that humans use their general intelligence to drive, since Humans can learn on the go and add almost any type of new knowledge and act upon it without even realizing it. The brain is not looking for anything in particular, just any type of information that would be relevant to driving.
Funny, I was not aware of any fully autonomous vehicle that actually works with or without it.
Generous helpings of lidar and radar to augment cameras is a crutch to help compensate for the lack of 500 million years of unsupervised learning that went into our visual cortex.
https://en.wikipedia.org/wiki/Tesla_Autopilot#Incidents
Just about every one of those fatalities can be summed up as "Tesla ignores large stationary object directly ahead". Lidar would have detected all of those objects and most likely prevented every one of those accidents. I think Tesla currently has the best vision and radar only system out there, so either the state of the art doesn't quite cut it yet without lidar, or there's a ML engineer at Tesla that really really hates fire trucks.
It is noticing stationary objects, because sometimes it breaks when the car approaches a bridge going overhead, which is also bad when the car too close behind doesn't.
You have to ignore some objects in front of you (even ones heading directly towards you) because you're going round corners, so it's never cut and dried
Yes and no. It's also just nice to have more different kinds of data available. Just because humans don't have laser eyes doesn't mean we can't try to do better than that.
Depth from stereo is not ML, but yes.
Tesla in particular, and others are moving towards getting depth from ML. Yes, you can do dumb coincidence-finding, but there's a lot of corner cases (leafy objects, specular reflections, etc) that screw this up. Humans don't just use coincidence finding, but use all kinds of other clues (size, texture gradients, monocular parallax, shadows, linear perspective, attenuation from haze, etc) to infer depth.
The human neural network is of course fine tuned to this and my father who is blind on one eye will actually move his head sideways a bit to judge distance when he is driving :)
Maybe not just, but we do use it, except we do it by converging our eyes for _actual_ coincidence within a small focal center instead of ambiguous coincidence everywhere in the field of view like CV stereo typically tries to do. Gimbaled cameras that (metaphorically speaking, of course) "look" where drivers are supposed to look with an attention model instead of being statically coplanar could do this too.
Only at very close range. That's convergence, and mostly a 5m and less phenomenon (not so relevant at driving ranges).
not amazingly well, though, because there are millions of victims each year
SDCs are supposed to be better than us, so more sensors make sense.
Also unsafe lefts and ran/stale lights. That's probably most all of them between our comments.
Radar: It can accurately measure Distance and Velocity information of objects around ego vehicle and also can track objects. It works well in all weather conditions (Day/night/rain/fog etc). Lidar: Good distance measurement, Rich in data (3D Point cloud) for ML, OK in doing classification (pedestrain, bicyclist etc) even in nights. But expensive sensor.
They're just building in the equivalent of putting in your sunglasses or looking off to the side when the sun is in your eyes or readjusting your seat position so it isn't hitting your eyes.
A system which costs too much to widely deploy and maintain will never see widespread deployment.
It's not, in theory. It is, in practice.
> > Is it solely because computers aren't yet able to extract as much information from video as humans can from sight?
Yes. Currently, you cannot, for any amount of money, buy a camera that has equivalent visual acuity to the human eye.
Your eyes can operate - and operate well - in an incredibly broad array of lighting conditions. No single camera exists that can currently do that.
Dynamic range is a challenge for cameras, but we have high-dynamic-range imaging nowadays (both in software, i.e. exposure combining, and hardware). I don’t think the eye is significantly superior here.
Low-light used to be a strength of the human visual system, but I think modern computational imaging systems have caught up (plus, human night vision was really never that good compared with other animals).
So in short: I dispute the idea that cameras don’t have the acuity of the human visual system, nowadays. I’d like to know in which aspects you believe human eyes to still be superior, from an optical point of view. Obviously - the human brain and visual cortex is something that computers are nowhere close to.
Dynamic range. Your eyes can pick up subtle details in a scene with very bright lights, and very dark shadows.
No video camera currently exists that can take a good video from inside a city, at night, with a starry night sky. Either the stars, or the lights, or the shadows are going to look like crap. Your eyes can trivially handle such a problem.
And even if Google has those extremely good filters, there is an entire different order of magnitude of testing until they can be sure. They are probably jut trying to get something out of the lab by using reliable sensors, and letting optimizations for later.
Human eyes and brains are completely different from solid state electronics with different constraints and advantages.
He's been a long time anti lidar proponent because of the costs involved and the aesthetics. He's also betting that the amount of data Tesla receives from its customers, and the neural net they have can achieve autonomous driving with its current hardware stack.
“In my view, it’s a crutch that will drive companies to a local maximum that they will find very hard to get out of,” Musk said. He added, “Perhaps I am wrong, and I will look like a fool. But I am quite certain that I am not.
"Despite being a fancy and expensive technology, LiDAR provides surprisingly little advantage over a combination of cameras and radar. Radar, for example, is much better in the rain and other limited visibility scenarios, because it is based on radio waves rather than light waves. Radio can penetrate through some objects and bounce back from others, thereby “seeing” the environment along a different dimension."
an excerpt from a quora answer on why the Tesla stack could be better than Lidar: https://www.quora.com/Why-dont-Tesla-cars-use-Lidar-like-mos...
Here's a video of Tesla's autopilot perceives its environment : https://www.youtube.com/watch?v=fKXztwtXaGo and here's a video of how Waymo perceives its environment: https://www.youtube.com/watch?v=OopTOjnD3qY
The question is if the Lidar adds incremental value or exponential value, and I think it's just incremental by looking at those videos.
This is biased information as there are way more Tesla's out there compared to Waymo's.
Waymo has had accidents too https://www.wired.com/story/waymo-crash-self-driving-google-...
I often see oncoming cars rushing to make a left turn past when their arrow has turned yellow and red, and so I know not to enter an intersection even though my light has turned green.
Autonomous systems in theory should be better than humans at this because they can track all surrounding objects and trajectories, not just ones they are looking at with one set of eyes.
I think the accident rate is kind of a meaningless stats. You can have a low accident rate by carefully controlling the conditions under which you test. Not many accidents means they aren’t pushing the envelope. That’s probably a good thing on public roads. It’s also why the system is not available for general public use except under extremely controlled routes and close (remote) supervision.
I think it’s great we have (at least) two mega-companies in a race trying different approaches to reach a solution. There are good points for and against both approaches. This is what makes life interesting, you can’t just run the numbers to predict the future.
Though I ended the answer at the end asking whether Lidar adds incremental or exponential value? Do you think it adds exponential value ?
I'm not an autonomous car engineer so don't understand the nuances but from whatever basic information I've read it doesn't seem like Lidar's add exponential value.
edit: Just to add, 5 million Waymo miles have been driven with a driver that controls the car and Waymo has had an accident too - https://www.wired.com/story/waymo-crash-self-driving-google-...
Also Waymo has way lesser cars on road than Tesla
The usual metric for self-driving car success is "disengagements per mile", ie how frequently a driver needs to intervene to avoid a crash. From my anecdotal readings of Tesla Autopilot reviews, it's on the order of 0.1 per mile. For Waymo and Cruise, it's on the order of 0.01 per THOUSAND miles. That's a very different definition of "driver" than the one that Tesla Autopilot requires.
I don't know the total number of miles on all Teslas on Autopilot, but it has had much more than one accident.
EDIT: and that Waymo crash was not a self-driving error; it was T-boned by a human-driven car running a red light.
The reason I assumed it works is that lidar on the article above seems more like a redundancy. Because their camera system have the short range covered and radar has the long range covered. Lidar seems to augment over it.
Though the order of disengagement is a great stat, that definitely shows how much better waymo is compared to Tesla
The usual solution is actually lidar + optical; lidar gives much better spatial resolution than radar, which is why it's been the standard going back to the DARPA challenge. You really want to have good spatial resolution in order to distinguish e.g. bikers and tail-lights and road signs for your optical systems, which radar typically isn't good enough for; that's the point of that qualifier in "imaging radar". Still probably worse performance (i.e. time and spatial resolution) than lidar, but better range and weather resistance.
(The previous generation of Waymo cars already had one lidar on top; the radar and the close-range lidars are the new additions.)
Miles per disengagement I’m sure is not too high. That would be a good metric to have. But total miles is still ~2 billion.
I can't help thinking that, whatever the merits of Lidar, Musk is boxed in, because Tesla has sold hundreds (tens?) of thousands of "self driving packages" for cars not equipped with Lidar, so changing course would not just mean raising prices on new cars, but retrofitting large numbers of existing cars at a ruinous cost.
From the systems I've worked with it's usually AND and not OR, you use both a Lidar and a Radar. The Radar images I've seen were quite lousy and are not 100% interference prone.
I'm not sure that unsubstantiated claims from Elon Musk are actual evidence that lidar isn't necessary.
... but we do want the SDC to do better, and there are failure modes that human perception is also vulnerable to generally. In addition to closing the gap faster on solving the problem without a copy of the human perception wetware, the LIDAR signal might also improve on those perception error states and be worth keeping in the design even if it could be done with cameras alone (or camera + radar + ultrasonic).
However, computers don't have human brains, and AI doesn't provide _anything like them_.
Some day, someone will make that bet and be right. I haven't put my money on this team and this project. ;)
Using a two camera input, the best we could hope to get was a depth map. Typically, it is reliable for at most one or two reliable levels of depth, useful for foreground/background separation. Maybe also a crude depth map, useful for fog. With Lidar we could all that plus normal information (i.e. face orientation).
However, I concede your point. The core difference is one of accuracy.
A fully autonomous driving system that was only as good as a human would not gain much traction and the company would be on the hook for some serious legal damages if it became popular.
But arguably very little progress (none?) in technology, automation, and industrialization has been through dogmatic replication of biological systems.
Planes don't fly like birds...
We don't have a working driverless car with nice easy LIDAR data - it would be really silly to try and make one the hard way first, using only vision data (yes I know about Telsa).
Of course, why would you not avail yourself of more sensory information if you can, provided the cost is not overly burdensome.
The only thing a camera can offer is input for pattern recognition. Lidar/Radar offers context.
I'm not suggesting that all you need to fool computer assisted driving is a picture of an empty road taped to the front of the camera; but if the computer can't tell the difference between an optical illusion and reality then I think they need more data inputs.
We don't need computers to be able to handle white-out blizzard driving conditions, but that 'rare' occurrence perfectly illustrates the singular limitation of a camera-based pattern-recognition system.