OTOH, no one cares what the nonlinearity actually is so long as the net trains, and there's a lot of effort to keep layer inputs in a neighborhood near zero (batch normalization), so polynomial explosion may not be such an issue.
Feel like I would like a more serious comparison of their model's results with best of class NNs. I'm suspicious that their NN character detector was basically failing...