The doomed part is because if companies are spending all of their energy on creating neural nets around lidar then they'll reach a local maximum where they never begin to tackle the much more difficult problem truly needed for self-driving.
The doomed part is because if companies are spending all of their energy on creating neural nets around lidar then they'll reach a local maximum where they never begin to tackle the much more difficult problem truly needed for self-driving.
The first time Tesla runs over a kid chasing a ball into the street because it couldn't see him between the cars, this will be readily apparent.
Seems to me that Tesla is in the business of selling cars, other self driving companies are interested in AV for ride sharing or trucking. The latter have different requirements for styling and costs and the consumer case, so Musk has several limitations on the sensor suite he includes in a Tesla.
What he's doing is trying to argue a $5k system with cheap cameras and crappy radar coverage is all that is needed, because a full no-blind-spot multi-spectrum system would both cost too much AND likely make the car look ugly.
Two people have already been killed, and several injured, by Tesla autopilot due to blind spots.
Will it be apparent how fundamentally problematic this is when a human runs over a kid chasing a ball into the street because it couldn't see him between the cars?
How many people have been killed by human drivers due to blind spots?
Musk's argument is more that cameras should be sufficient because humans can drive using only two eyes to perceive the driving environment. He always neglects to mention that humans do this with a combination of sight and a brain capable of general intelligence. I'm sure it's true that if Tesla invents AGI, self-driving with just cameras will become tractable. But "real hard problem to solve" doesn't begin to capture the difficulty.
In reality, since no one has yet invented a self-driving computer, it's impossible to say what components are necessary or even whether there may be more than one way to skin the cat. But one source we should probably take with a grain of salt on this issue is those (like Musk) with an intense commercial interest in one perspective.
"So this guy says a machine can outplay chess with just a CPU, some memory and a bit of code. He neglects to mention that humans have a brain capable of general intelligence".
Year later, computer beats human in chess.
"So this guy says you can teach a computer to play Go just by unsupervised training of neural networks. He neglects to mention that humans have a brain capable of general intelligence".
Year later, computer beats human in Go.
"So this guys says you can program neural network to play computer games competitively using vision and deep learning. He neglects to mention that humans have a brain capable of general intelligence".
We don't need AGI for self driving.
The difficulty of self driving is probably less than a dog walking on the street.
No, a dog doesn't steer a car, because he doesn't have hands, but he's performing the same vision and planning tasks like a human (or AI) driving a car.
He knows where he is, he knows where he wants to go and he uses vision and his non-AGI brain to plan a path to get there while also avoiding dynamic, unexpected obstacles.
It seems like maybe not the best reasoning to say “humans, do it this way, so that’s how my robot should do it.”
I don't know why Tesla autopilot keeps missing obvious impervious occlusive surfaces, and detecting obvious impervious occlusive surfaces is what Lidar excels at, it's kinda making the point for the other side.
It doesn't look like it will be widely available for a few years at least though.
Well, no.
RGB is really useful. But actually, you can get pretty good object recognition from a point cloud alone. I mean its better to have RGB as well. but infra-red works just as well.
The problem that appears to escape a lot of the commentary is latency. Sure, you can have a rudimentary stereo camera setup and get _some_ depth information reasonably fast. But it won't be good enough to tell you if that blob that's 100m out is stationary or moving towards you.
Lidar gives you high resolution long range 3d point cloud at 30hz (or faster). The best most reliable depth from monocular/stereo will have a latency of at least 150ms and will be a tiny resolution.
The chances are that we will have sub $100 CCD based lidar before we have low noise/low latency/full resolution depth from monocular/stereo cameras.
The other big issue is that to get decent high res depth from deep learning, you need to have decent segmentation. Segmentation comes for free with lidar (assuming you impose rgb over it.)
> spending all of their energy on creating neural nets around lidar then they'll reach a local maximum where they never begin to tackle the much more difficult problem truly needed for self-driving.
This does not make all that much sense. You don't just train on lidar, you feed in steering, acceleration, braking, gears, signs, radar, pretty much everything.
The other important thing to note is that tesla's stuff is still level 2. volvo, BMW, and a few truck companies are all at least level 4. We are celebrating a "genius" who has yet to actually release a system that does what he claims it should.
They admitted their own models are far from perfect and will likely never be. The concerning one in particular was the "is this a large object" model which initially failed to identity objects such as car-carrying trucks, cranes etc.
With Lidar you can certain at least that it will identity an obstacle.
Why can you only achieve that using visible light spectrum video?