The question of Camera vs LIDAR+Camera is a narrow technical question about how to construct a 3D scene. That's it. It says nothing about making sense of this 3D world for which you you have a 3D point cloud and it says nothing about how to actually navigate that world. Say you're driving down the road and there's a bit of construction, there's a guy holding SLOW/STOP sign directing traffic. LIDAR will tell you it's a hexagonal sign, but it can't tell you what it says, you need a camera to read the sign and tell you what it says. It doesn't tell you how to drive, how fast you should go, how much space to give the guy with the sign etc. Everything AV-related which is not constructing a 3D scene is actually the same across all AV stacks, which includes the hardest part - the actual driving itself.