Re: the book added in your edit. I have it right in front of me.
As I said before, we know a lot of facts. We know a lot about the spectral sensitivity of rods and cones, and the molecular mechanism that lets them turn photons into electrical impulses. We know a little bit about where the areas that process faces are and what visual features the neurons in them respond to. We’ve got pieces, but they’re not put together.
I would say that we understand vision when we can answer a question like “How do you find a friend in a crowd?”
You can start with “When you first met, light bounced off her face and isomerized some retinal from its 11-cis to all-trans form, which caused the bound opsin to change conformation into metarhodopsin II, which activated transducin, which....” Eventually, this cascade caused electrical activity that reaches cortex. A huge set of cortical areas process visual input, and these electrochemical signals flow through all of them. We can predict V1 neurons’ activity reasonably well, less so for the downstream neurons in V2, V4, or the temporal lobe areas. We have only the fuzziest ideas how those patterns are read out, tagged as important to remember, and moved into memory. You've only just met--and yet it gets worse.
To find her, you’ve got to retrieve those patterns from memory (no one knows how, but oscillations might be involved?), and use them to search in a way that’s robust against variations in the friend’s pose, position, rotation, illumination, and even dress style or age, many of which you have never seen before and will never see again. We know, for example, that some cells in IT are fairly robust against some moderate kinds of image changes. Some but not all, of this is done by circuits that look like a convNet. Whether this is a coincidence or not is debatable and how this convNet is trained is a total mystery—-it’s definitely determined by experience, but the feedback signals needed for vanilla backprop are missing.
As you scan the crowd, you’re only getting high-resolution data from a very small part of the visual field. This is (somehow) stitched together into a unified percept. You apply various heuristics—maybe your friend favors bright colors—to speed the search along. How you learn these, and how they’re mixed in with the input from your eyes is unknown, but it’s certainly reflected in your behavioural output: you'll find her faster if you successfully predict what she looks like, and you'll be much slower if you guess wrong. Perhaps you hear a familiar voice or smell the perfume you bought her. This, too, can help you find her, but how information is integrated across senses is unknown too.
Eventually, you find her. You plan a path across the plaza towards a cafe. We have a pretty good understanding of how this works in rats (3-7 Hz oscillations coordinate place cells and grid cells in the hippocampus). Those oscillations are really strong in rodents, but much weaker in monkeys and totally missing in bats, so it's not clear how this works in humans.
Now all you’ve got to do is open your mouth and order coffee....