One of the biggest challenges that automated systems face is that the acceptable failure rate for them is far below the acceptable failure rate for humans in the same role. To err is human...
The difficult part - when there is an error in these signals, or things shut down, autonomous cars will suffer much bigger problems than human driven cars.
Its these edge cases that are the problem. We already rely on such mechanisms for planes(information comes from both gorund control and on-flight radar). But a lot of care and resources are is required to get to $n 9's level of reliability.
Sounds costly to me. There are a lot more roads than airport runways. And then the big question is: Who is going to pay for it?
The situations you describe are rare - I've once had a diplomatic event that required weird rerouting and twice had cases where traffic was regulated by hand signals due to some crash on the road, but that means just a few cases over a whole lifetime. A system that can't solve these cases but recognizes them as unsolvable is a quite acceptable automated system if it can delegate control to a human inside or a remote dispatcher, which isn't that hard to do.
And this is just me driving (i.e. my car is parked 90% of the day). If you're talking about a self-driving Uber in D.C., one of the above events will happen on a daily basis.
http://www.masslive.com/news/index.ssf/2017/05/east_longmead...
Not that's not sufficient.
If ten people did that in a critical area during a high demand hour it would be a news story and there would be criminal charges depending on the details.
If you redefine "sufficient" to include stopping your car on the George Washington bridge because it's confused by a construction zone it still doesn't solve the backup you cause.
Of course, there's a very reasonable argument that e.g. level 3 automation might cause fewer accidents overall, even if it kills people when it has no idea what to do, but convincing Joe Public that such a car with such a known flaw is safe is another matter.
A human will spot a person wearing headphones and recognize that person has a low situational awareness. The computer doesn't come close to even having the optical resolution to do that if the AI was perfect - remember human vision is 570+ megapixels, even a 4K video stream is literally two orders of magnitude lower.
[Now think about the fact that if we built a camera capable of recording 400 megapixels, you'd currently need to schlep around a ~750 lbs 25 node cluster, consuming about 50 horsepower to feed it with electricity, just to be able to process the video stream at 25 fps. Moore's law aint' growing that fast these days, so matching the resolution of human vision is not a realistic option.]
Another example is kids. How does the AI recognize that the 5'1" 30-year-old woman has much better awareness and can be treated differently from the 5'2" 12-year-old boy? Humans can spot that difference even from behind.
How about recognizing an adult who is drunk? Or a blind person? Mourners at a funeral, or fans celebrating after a football game? Or a million other conditions that significantly affect pedestrian situational awareness that human drivers will instantly infer from context?
What will happen when kids figure out they can stop a driverless car on its way to collect its owner just by standing in the street in front of it? They'll have a lot of fun, for sure.
How about when carjackers figure out the same? That they can dress up like construction workers, stop the car in the street, tow it onto a flatbed with built-in RF jammer and head straight for their underground chop shop? There goes your cheaper insurance.
That all people are classified as drunk children wearing headphones with low situational awareness.
This seems to come from http://www.clarkvision.com/articles/eye-resolution.html
But that number is a calculation of the maximum resolving power of the human eye filled across a 120 degree field of view. The fovea is the only portion of the retina that actually attains that acuity and it encompasses roughly 2 degrees in the center of the retina.
There are roughly 120 million rod cells and 6 million cone cells in the retina. The rod cells for color vision and cone cells for low light. As each individual rod cell is primarily sensitive to one of red, green or blue they match fairly well to the rgb channels of a pixel. So the eye could be considered to provide data roughly equivalent to a 40 megapixel color image and grayscale 6 megapixel. So ~5 times a 4k image.
Edit: And even that actually over estimates the amount of data the brain is actually processing. A 4k 60 fps video is handled by 6Gbps and the human optic nerve only has roughly 8.75Mbps of bandwidth.