Next step is Video. Adding temporal dimension will emphasize extrapolating true 3d shapes of recognized objects.
It seems that this is more like the way that we learn to identify things. Then once we establish an understanding of a base class (big cat) we can apply that same model to new cats that we have never seen before with just a picture.