I’m not convinced that the camera-based approach is just a learning problem. As far as I know, cameras still have awful dynamic range compared to humans. Remember the Uber crash video where it looked like the victim was only visible for a couple seconds before impact? Cameras need to work a lot better to be suitable for autonomous driving.