LSTM's are clever and they were bleeding edge up to 2014, but once they were understood better as a bypass mechanism, attention, context vectors and averaging networks and causal convolution are starting to replace them.
When you put researchers or groups into order of importance, the work from this trio comes before others, including Schmidhuber. The Turing Award is given to those at the top and it's not inclusive.
If I could give Turing Award for someone in the field who has not received it yet, I would give it to Vladimir Vapnik.
So would I, but I don't think that will happen. He's is pretty much an antithesis to everything Deep Learning is about. Theory-first over trial-and-error, math over intuition, small datasets over big data, advances in understanding vs advances in results.
I haven't seen a single discussion of his paper on combining classifiers[1] anywhere on the web.
[1] - http://jmlr.csail.mit.edu/papers/volume17/16-137/16-137.pdf
"Please respond to the strongest plausible interpretation of what someone says, not a weaker one that's easier to criticize. Assume good faith."
As far as I'm aware, causal convolutions were used in WaveNet (and subsequent models) and a small number of NLP applications. Meanwhile, LSTM-based models are used in just about every NLP paper, and at least a baseline in the newer ones more dominated by Transformers.
But that's exactly the point of the Transformer model, with a paper aptly titled "Attention is all you need" [1]. And the Bert architecture, based in this idea, seems to be doing well. And they claim to be bery flexible, too[2].
Maybe that's what you meant with "unless you brutely search over hundreds of hyperparameters configs", but then again, isn't that what NNs are about anyway?
Out of curiosity, at the same level as any of the other three? I could see it from Hinton, maybe LeCunn, but Bengio?
Hinton/Lecun pioneered early BP research, which is significant and fundamental. Bengio has attention mechanism/GAN, and in early days greedy layer wise pretraining for RBM (which isn't something end up today, but pertaining is a thing back in pre 2010 era).
Not saying LSTM isn't significant, but his achievements aren't quite there with those 3, looking closer.
That's not true with Schmidhuber.
http://people.idsia.ch/~juergen/deep-learning-conspiracy.htm...
Note for people put off by the link: The conspiracy refers to the title of a paper by the Turing award winners, not Schmidhuber's.