If anything, knowing when to reliably ignore a sensor modality is the kind of intuition more associated with general AI.
A similar paradox occurs when trying to fuse multispectral imagery. You'd think early fusion of RGB and IR would be better since it gives the higher-resolution filters access to more data, but it does worse than late fusion. My understanding is that late fusion forces the network to "work harder" to solve object detection using IR only, and then once you've wrung what you can out, then you fuse with RGB detections.
Since radar is "one pixel" there's essentially only one object detector possible: object or nothing. If yes-object, fusion tries really hard to make sense of the RGB filters to figure out what partial detection looks like an object, which is almost always a false positive.
Waymo has its own custom developed 4D imaging radar and I imagine most of the other SDC players have their own versions.
[1] https://blog.waymo.com/2020/03/introducing-5th-generation-wa...
[2] https://aurora.tech/technology/driver
[3] https://www.mobileye.com/blog/radar-lidar-next-generation-ac...
Given this post is about radar usage in self driving and how Tesla/Karpathy touts their capabilities, I compare them to other SDC players, not other car manufacturers who are not in the game. And their SDC competitors are running away because they have the newer 4D radars.
Only if you don’t consider Waymo, Cruise, Mobileye and others as Tesla competitors.
+1 to the other commenter (grandparent comment).
Radar has a range resolution and an angular resolution. Andrej completely omitted this fundamental aspect in his talk.
In my humble opinion, Andrej needs a real Radar engineer in his group before making such a statement in his talk.
Nobody can be smart for all topics, which makes all humble also.
That sounds like a problem with the network's architecture and not the data itself
This makes some important assumptions, namely that Tesla built a lidar and radar perception pipelines and sensor fusion of equivalent quality to their competitors, and then decided they were unnecessary.
Given that their competitors have shown substantially better perception than Tesla, and that Tesla has a significant economic incentive to deliver autonomous driving on a sensor suite that already shipped years ago, I find that difficult to believe. Did Tesla build good enough perception to dismiss lidar and radar purely on their merits. Unlikely I think. Did an intern build a student-quality lidar pipeline that "proved" Elon's camera-first approach is the right one? More likely.
I am actually on sensor-fusion side, and think a transformer can merge everything and generate a coherent world-view. But this is a hard problem and people side-step them after the evaluation shouldn't be dismissed blatantly.
For one:
> It gives you a more robust view of the world than using a single sensor modality.
How can we evaluate that correctly?
In effect, they solved the disagreement problem by sticking their head in the sand and pretending their sensors are never wrong.
This is actually my main criticism of Tesla's approach. They don't seem to have enough cameras to do the job well, and its showing in a lot of actual system limitations.