I consider the paper “ImageNet-trained CNNs are biased towards texture; increasing shape bias improves accuracy and robustness” essentially mandatory follow-up for anyone interested in this. It goes further with its conclusions and introduces a potential solution. By introducing the ‘Stylized-ImageNet’ dataset, they found that they can force the network to learn shape data, improving scores even on standard ImageNet and preventing this BagNet trick from working.
https://openreview.net/forum?id=Bygh9j09KX