Tesla auto-pilot keeps confusing moon with traffic light then slowing down
reddit.com
reddit.com
All Hail Elon.
That said, the computer vision is still far from human level. It lacks a common sense around how the world works.
You'd be able to tell there's no object anywhere near the car and disregard the light.
Lidar would be able to tell you if the object should even be evaluated.
A light hanging from a cable would still get a lidar ping back. The moon would be a void with no lidar ping.
Maybe is not acceptable (and blind in the rain/fog), so you must rely on vision to make the decision – which is why Tesla goes straight to the heart of the problem.
They haven't solved it yet, obviously, but neither has Waymo or anyone else. Tesla is probably collecting more images of 'moon or traffic light?' images than the rest of the industry combined already.
In the above video Andrej Karpathy claims that LIDAR required previously compiled maps, the creation of which is not scalable. Do you disagree with that point of view?
Without the maps, you'd need to rely pretty heavily on vision to do the right thing.
What LIDAR gives you is a very accurate point cloud to work from. You can get a point cloud from vision (and indeed, that's what telsa is doing) but it won't be near as accurate as LIDAR is by default.
Their automotive gross margin is 28% (Toyota is at 17.76%). If one automaker could afford LIDAR, it's Tesla.
That said: - way cheaper - more reliable - less power hungry - simpler and easier to debug.
An autocar seeing something that looks like a red light and breaking (or breaking for a bit and then continuing after deciding it was not a red light) is going to be acceptable to people because that is how they behave. Phantom breaking due to some radar or lidar input that the human can't perceive is going to be interpreted very negatively. I imaging auditory might be added at some point, if it is not already being used.
1) They have hundreds of millions of cars sold already with cameras that they can up-sell this to. That means more revenue. But more importantly, more training data. Massive amounts of it. Training data is the real value here. Radar/lidar/etc. might be able to detect better what is obvious to a human just looking at a thing. But given enough training data, a machine learning can probably replicate that capability so you don't actually need the fancy sensors. Adding new sensors to the mix would set them back quite a while on that front.
2) Simplicity. More sensors means more complexity integrating all the signals and gathering the right training data. More failure modes, etc. It probably also means more compute power needed to process all that data. More complex testing, etc. Scaling by keeping the sensor platform simple is a good move here.
3) The hard part of autonomous driving is actually interpreting visual signals in complex or unusual/rare situations. Roads are designed for humans with eyes. Lidar sees a blob, radar detects a pole, a camera sees a traffic sign, road markings that mean something, etc. It's a much richer signal. All the important stuff on roads is clearly visible. So, cameras are far more important for this than lidar/radar. Those are really great for avoiding crashing into things. Not so much for interpreting and classifying those things. And Tesla seems to be doing pretty OK with not crashing into things. Mostly, the amusing edge cases have to do with misinterpreting visual signals for which radar and lidar are probably not that relevant.
It's an interesting approach that they clearly believe that they can make work. It does not actually stop them from later adding more hardware to enhance things if they decide those things are needed. But it's quite interesting how far they are getting with just cameras.
The other two sound like good reasons though
Obstacle detection with lidar/radar is only interesting if you assume that object avoidance is a problem with camera based obstacle detection right now. There are lots of incidents with Teslas but I don't recall them running over pedestrians or crashing into vehicles a lot. Mostly incidents are about misinterpreting visual signals; not about crashing into stuff. If anything, their safety statistics are pretty good when autopilot is on. The cars still do dangerous/illegal/misguided things due to misinterpreting of traffic situations but it's then smart enough to get the driver out of trouble before bad things happen.
http://taihendaro.cynic.net/2010/01/moon-as-soviet-missle-at...
Tesla’s vision/ml systems are amazing. I would love to learn more about how unit testing for this type of error is done. Without some intermediate semantic representation, I don't see how these large, multi-head, end-to-end systems can isolate and regression? System tests are maybe possible, but it's unclear how well a system test would generalize to related, but unseen cases.
More importantly, they're the only ones accessible to the average consumer. Waymo/Zoox/Daimler all have equally if not more impressive systems.
A real issue with tesla is that they want to be vision-only, which is going to make getting to level 5 first almost impossible.
BTW - knowing where the moon is happens to be an extremely solved problem.
It's true... but I don't think that's how Tesla would solve it. Their goal is to create a neural network "driver" which can drive in any place even if it has never seen it before. They'd rather teach their neural network that the moon and stoplights are not to be confused visually. Thought I suppose in searching for training examples they could use the known position of the moon for approximate labeling.
Isn't that impossible considering these networks need training and therefore have seen everything before?
But with the Tesla, it is learning to drive in general. It does not need an HD map of a fork in the road to understand how to navigate it. Just as a person who learned to drive in California will have little trouble driving in Florida, a neural network that has learned to drive on a million intersections will be pretty good at navigating most intersections. Especially because the corner cases will stand out and become integrated in to training. So it may see many intersections, but it will generally know what to do with one even if it has never seen it before.
Though I would suspect that the competitors are perhaps using HD maps to jumpstart a system that long term would behave more like the Tesla one. Mapping every road is a lot to ask.
One actually really important feature of these systems is how they handle failure. If the car gets confused, how does it handle it?
But the big thing is that their autonomy computer can be programmed to look for odd scenarios and send them back home. Tesla uses their fleet of hundreds of thousands of cars to collect edge cases like this, and then they have a kind of compartmentalized neural network system that breaks apart disparate tasks. With their collected examples they can create unit tests to ensure that the moon stops activating the stoplight detector. Once trained, the unit tests presumably help ensure they don't end up with future regressions.
So basically every time you see a Tesla do a weird thing, there is a good chance it will stop doing it soon enough. At least if it's hitting hacker news.
Permanent yield signals are often only one flashing yellow light. Crosswalk signals are often only flashing yellow lights when active. Temporary construction barrier signals are often only yellow flashing lights. Fire station signals often have only red and yellow lights, no green. Even when all three lights are present, they may also be oriented horizontally, or in triangular shapes.
Which isn't to say there isn't more context to learn from, but just about the only true unifying trait among all these indicators that you should perhaps pay attention and slow down is a bit of yellow light. It need not even be circular: yellow arrows are far from uncommon, including "straight ahead" yellow arrows for intersections where turns are forbidden.
Edit: this also reminds me of those Road Runner episodes where Road Runner paints a rock wall like it's the continuation of the road so that Coyote can smack his head on it.
EDIT: Though even that isn't a traffic light, and not the other way around.
A couple decades later, I'm now noticing some traffic lights in the US starting to gain this feature. I'm sure it will help with machines trying to parse the traffic light state... and also humans.
Not all the time. I've seen reddish/orangeish moons, especially if it's something near the horizon near sunset or sunrise. It's also reddish during lunar eclipses.