I've always wondered... if Lidar + Cameras is always making the right decision, you should theoretically be able to take the output of the Lidar + Cameras model and use it as training data for a Camera only model.
So...nowhere?
Why should you be able to do that exactly? Human vision is frequently tricked by it's lack of depth data.