I'd be interested in seeing if there were any "attributes" that they could not easily explain.
I suppose that a well trained model shouldn't have too many of these, but it would be good for things like detecting intentional manipulation of training data.
This also makes it way easier to fool classifiers by exploiting known high-impact attributes.