WaveNet, if I remember correctly, has 1-from-256 encoding of input features. And 1-from-256 encoding of output features.
It is extremely sparse.
If you look at language modeling, then things there are even sparsier - typical neural language model has 1-from-several-hundredths-of-thousands for full language (for Russian, for example, it is in range of 700K..1.2M words and it is much worse for Finnish and German) and 1-from-couple-of-tens-of-thousands for byte pair encoded language (most languages have encoding that reduced token count to about 16K distinct tokens, see [1] for such an example).
[1] https://bellard.org/nncp/
The image classification task also has sparcity at the output and, if you implement it as RNN, a sparsity at input (1-from-256 encoding of intensities).
Heck, you can engineer you features to be sparse if you want to.
I also think that this paper is an example of "if you do not compute you do not have to pay for it", just like in GNU grep case [2].
[2] https://lists.freebsd.org/pipermail/freebsd-current/2010-Aug...
Given all that I think it is a paper about combination of very clever things which give excellent results in a synergy.