> We find that the features that are learned are largely universal between different models, so the lessons learned by studying the features in one model may generalize to others.
Hm. I wish they'd said more about that. Does that mean they found the same feature recognizers when training with the same training set? Or what? This tells us something, but what does it tell us?