Adding more sensors slows his team now more than it improves system performance
I'll take his word on this. It is a lot of work to incorporate multiple sensors.
All necessary information is already in the pixel-space.
I hate to disagree with someone as distinguished as Karpathy, but this is simply not what I have observed from all of that data that we have access to. Given my knowledge of the various stacks deployed today, I would never ever ever get into a vehicle using a vision only stack and expect it to perform in some of the challenging environments encountered during testing.
On the other hand, RGB data does have that information, we use it everyday to avoid obstacles, even under foggy and rainy conditions (I'm no LIDAR expert but I know it sucks in rainy conditions)
I am not saying I support a vision only stack, but all I am saying is it is certainly possible to deploy a vision only stack in the future.
The fact that (most) humans manage to drive around safely and successfully in current roads proves that the information needed exists in the pixel-space (not just current image, but say current + history). We don't yet have stacks that can successfully map everything needed from this information but I don't think Dr. Karpathy ever claimed that.
(I am not a principal engineer but a mere PhD student who argues daily with people on how RGB information is underappreciated and under utilized)
I also agree that most humans manage to drive in challenging conditions, but their margins for error become slimmer and slimmer. I personally want my autonomous robot vehicle to be way more efficient and safer than the best human operator and also able to deal with conditions that any sane human would pull to the side of the road when encountering.
In some way, I am against the philosophy of using HD maps + LIDAR data for highly accurate localization which most companies seem to be using these days. I believe that this approach is inherently brittle and is an 'easy way out' to the hard localization problem. I think more resources should be put into developing more natural, no HD map dependency techniques.
PS: It is my understanding that most of the major players were using HD maps, not sure if it is still true.
Can you elaborate on this? I've always felt like the margins of error are getting wider because the automotive tech (particularly safety features) are so vastly improved. I doubt people would be able to text and drive as much, for example, if they were driving a 1950s era Willys jeep just because it requires so much more attention to keep on the road by comparison to modern vehicles.
But that doesn't mean that it translates to a car.
We constantly move our 576MP resolution eyes in multiple orientations in order to visualise a scene and focus on the most important areas. Cars have fixed, low-quality cameras.
We then interpret this data using the most advanced pattern recognition system the world has ever seen that is trained for at least 20+ years to fully comprehend the behaviour of everything this planet has to offer. Cars don't have anything close to this.
Humans certainty have a stronger and general prior to make sense out of the information, and that's exactly why I left it as a possibility. Cars don't * yet * have anything close to it, just like they didn't have a way to accurate detect objects a few years ago and just like they didn't have a way to capture RGB information a few decades ago.
I am an optimistic guy, and I certainly believe in the power of learning at scale.
Actually our eyes are more like 8MP: https://www.picturecorrect.com/what-is-the-resolution-of-the...
Perhaps higher synthetic resolution from moving our eyes about, or perhaps that is meaningless.
https://en.wikipedia.org/wiki/Preventable_causes_of_death#Am...
https://capitolfax.com/2021/01/26/aaa-wants-us-to-stop-calli...
Doesn’t mean it’s better or easier