Also, I don't think a "restart from scratch" is required when moving away from Lidar - I don't have any hard evidence to make that claim and am happy to change my mind, but I haven't found any argument for doing so yet.
Also, I don't think a "restart from scratch" is required when moving away from Lidar - I don't have any hard evidence to make that claim and am happy to change my mind, but I haven't found any argument for doing so yet.
I’m not convinced that the camera-based approach is just a learning problem. As far as I know, cameras still have awful dynamic range compared to humans. Remember the Uber crash video where it looked like the victim was only visible for a couple seconds before impact? Cameras need to work a lot better to be suitable for autonomous driving.
So, in a direct comparison where cameras are allowed to use a similar set of tricks tricks including a fair amount of post processing they actually win. Aka, Retna vs CCD the CCD wins, iris vs iris camera wins etc.
As for dynamic range: HDR and the like can easily take care of that. The proliferation of phone cameras has thankfully taken care of this. I don’t think cameras are even remotely a bottleneck here.
My guess is that LIDAR is superior in that scenario, but I've never tried a Sony A7M3 either.
I always had the feeling they softened the image. Or at the least it was camera with low quality/settings giving a false impression of what the view was. If you pause the video right before impact the detail is still much poorer than you would expect.
https://www.reddit.com/r/SelfDrivingCars/comments/869olw/hdr...
It's more likely that the sensors in the Uber setup are set for fast processing, i.e., less light per frame, rather than for visual fidelity, i.e., more light per frame. After all, faster processing is more important to a vehicle traveling 60 mph than sharp borders. If the shutter time is set low enough, you would get something similar to what Uber released even on a DSLR.
Also, as others have pointed out the Uber video was from a dashcam. The point of the dashcam was to record events in the event of an accident, not be a high-fidelity visual record.
A single image is always going to be a poor way of judging the light levels, because the exposure (shutter speed, aperture, ISO) can make it look drastically different.