Most people don't realize that computer vision research is done using tiny sets of tiny images. Think hundreds to thousands of images that are 32x32 grayscale or black and white. Google did 10 million 200x200 color images.
For a conservative comparison human vision can be thought of as two 15 megapixel cameras capturing at 30 frames per second. (The real numbers are higher and much more complicated.)
I suspect that given this amount and quality of input data most machine learning algorithms would have produced better than state of the art results.
Now the bad. Their model totally ignored time.
Most computer vision researchers work with static images. In biology there is no such thing as a static image. Everything moves. You move, your eyes move, the environment moves. The brain, as an exception, can understand static images, but the rule is motion.
A majority of features in biological vision are tied to motion. Light/dark change cells, moving edge change cells. A huge portion of the information you derive from sight comes from the temporal context of a given moment. The change from moment A to moment B is directly encoded and learned and is often more important than the state at A or the state at B.
I'm not faulting them, in fact, I am extremely happy with their results, but I want to point out that these results can and will get much better in the near future as the input size increases and people start to integrate temporal features into their models.