Even a human brain has to train for ~4-5 months to become interested in shapes (https://en.wikipedia.org/wiki/Infant_visual_development). Once the human brain has been trained for these basic shapes for a while, it is able to quickly break down a new class (i.e. a cat) and recognize similar patterns in other images. This is something that is very similar to the way that training a deep NN works.
Also don't forget that the current NN are being trained mainly for photos, not moving images. A brain may recognize a cat by its tail-wagging or fur movements, which is a dimension that is completely missing from still images.
Also check out this similar post and discussion: https://news.ycombinator.com/item?id=9247851