The problem with Musk's argument is that there isn't any indication so far that the machine learning has tricks up its sleeve that are sufficient resolve all of the ambiguous situations drivers come across. Having direct, accurate measurements of the environment serves as a ground truth that you don't have with just cameras. Humans don't have LIDAR, but we do have a different ground truth: we have a mental model of the world, and can reason about it. It's not just that we know things like 'fast cars take time to stop' or 'he looked at me so is probably aware of me', but rather being able to synthesize based on really complex behaviors we've never seen before, and in many cases,
direct communication. All that is to say, there is a lot more to human drivers than their eyes, and for the moment there is nothing showing that we're on the cusp of developing real AI that can address those concerns. I don't think it's just a matter of labeling more data and GPUing bigger models.
So while Musk might be right if we get such AI, it's not clear that it's about to arrive, and even if it does, we can still use LIDAR. In fact chances are that if you have such great AI, additional sensor inputs will turn it from 'Level 5 Autonomous Pilot' to 'Ultra-super human pilot who can preemtively react to situations you didn't even realize were possible in the moment.'.
Of course there are questions at hand:
1. Is Elon right and I'm wrong? Maybe strong(er) AI is right around the corner.
2. Is current LIDAR tech good enough to reliably provide the ground-truth I spoke about?