Lot of good comments here - CV is heading in that direction, but in addition to the other mentioned problems, the needed data bandwidth rises exponentially - my own take is that we go broad in some ML work where we should go vertical (I believe that the same thing that makes you jump when you see a (non) snake is what your brain is using for image fusion feature detection. However, the real problem not mentioned is parallax - we understand some bits of it, but our brain does a lot of high level manipulation of the data - in the real world, it's two completely separate incidence angles for 2+ visual sensors seeing a point at a given angular relationship, but the same point is not the same thing - take an index card, paint one side red and the other green, and look at it edge on far enough away both eyes see each side, and its close enough for parallax - sometimes produces interesting illusions :-) When you get to complex 3D structures and parallax, the problem gets a lot nastier, because after all the work to get the data you want, you are then forced to throw offending bits of it away and start making educated guesses - my own take is "You do not see with your eyes, you see with your mind" :-)