Still... I'm not very knowledgeable about AI, but the part of this that describes biological mechanisms really feels handwavy to me.
What we have in our heads is not a cleaned up version of the input. There are no pixels or symbol strings in the head.
All we have in our heads is big activity vectors that cause more big activity vectors.
Big activity vectors means nothing. What is the structure of these vectors? Even if the brain's image processing turns out to have a lot of random complications and exceptions that fit the data your eyes have seen before (even if these innumerable complications are necessary for it to work well), surely some god with perfect understanding of it could still summarize it as generally following such-and-such algorithms. Even if you don't know what the neurons in your model are for, I bet most of them are in fact "for" something comprehensible by mere humans, and understanding that would make it easier to design good starting conditions for models. Not that I can prove it...
(Note: I suspect that the author of this presentation has a more nuanced view than what I'm complaining about - or maybe I'm just misinterpreting it. This comment is just intended to start discussion about what's written on the slide.)