Not just 3D shapes, but understand actions as they develop in time with recurrent neural networks.
It seems that this is more like the way that we learn to identify things. Then once we establish an understanding of a base class (big cat) we can apply that same model to new cats that we have never seen before with just a picture.