Ignoring the fact that the title of this post is misleading and also not used by the linked article, I was interested to read that MaxNorm was _so_ effective. In my experience it is rarely used in the state-of-the-art ConvNets, though maybe that is because they are trained on large ImageNet datasets where overfitting is less of an issue? Weight decay/L2 norm seems almost ubiquitous is comparison.
Have other HN readers found MaxNorm to be that useful? Am I missing out?