> People drive reasonably well using vision primarily
This is not accurate however. Other important senses in use include proprioceptive, hearing and tactile feedback from wheels. In addition to vision and the improved dynamic range of eyes, there is the important fact that human vision integrates a world model into expectations. Human vision also models time and motion which help manage where to focus attention. Humans can additionally predict other agents and other things about the world based on intuitive physics. This is why they can get on without the huge array of sensors and cars cannot. Humans make up for the lack of sensors by being able to use the poor quality data more effectively.
To put this in perspective, 8.75 megabits / second is estimated to pass through the human retina but only on the order of a 100 bits is estimated to reach conscious attention.
> Computer learning networks can classify imagery at least as accurately as humans and sometimes more so.
This is true but only in a limited sense. For example, when I put in the image on the right (of a car in a swimming pool) from http://icml.cc/2015/invited/LeonBottouICML2015.pdf#page=58 (which you should read and find the talk of but) in ResNet I get as top results:
0.2947; screen, CRT screen
golfcart, golf cart
boathouse
amphibian, amphibious vehicle
For LeNet it's:
0.5422; amphibian, amphibious vehicle
jeep, landrover
wreck
speedboat
The key difference is learning in animals occurs by breaking things down in terms of modular concepts, so even when things are not recognized new things can be labeled as a composition of smaller nearby concepts. Machines cannot yet do this well at all and certainly not as flexibly. Things as lighting and shading do not move animals as much in the concept space.
> The execution strategy appears to be to run classification and command prediction all the time, and while the human is in control consider it supervised learning.
This strategy will not learn from accidents because the signal there will be far from optimal usually.