Definitely don't read it as a history. It's just a lie. Schmidhuber is laying claim to a lot of things he didn't do. And is taking anything that kind of relates in words to modern techniques and claiming that he invented the technique. Even if practically his papers have nothing to do with what the words mean today and had no influence on the field.
These are basically the only outliers who claim that automatic differentiation was invented by Linnainmaa alone. Many people invented AD at the same time, and Linnainmaa was not the first. Simply naming one person is a huge disservice to the community and shows that this is just propaganda, as much of Schmidhuber's stuff is.
1. First Very Deep NNs -> This is false. Schmidhuber did not create the first deep network in 1991. This dates back at least to Fukushima for the theory and LeCun in 1989 practically.
2. Compressing / Distilling one NN into Another. Lots of people did this before 1991.
3. The Fundamental Deep Learning Problem: Vanishing / Exploding Gradients. They did publish an analysis of this, that's true.
4. Long Short-Term Memory (LSTM) Recurrent Networks. No, this was 1997.
5. Artificial Curiosity Through Adversarial Generative NNs. Absolutely not. Andrew G Barto, Richard S Sutton, and Charles W Anderson. Neuronlike Adaptive Elements That Can Solve Difficult Learning Control
Problems. IEEE Trans. on Systems, Man, and Cybernetics, (5):834–
846, 1983
6. Artificial Curiosity Through NNs That Maximize Learning Progress (1991) I have nothing to say to this. This isn't something that worked back in 1991 and it's not something that works today.
7. Adversarial Networks for Unsupervised Data Modeling (1991) This isn't the same idea as GANs. The idea as presented in the paper doesn't work.
8. End-To-End-Differentiable Fast Weights: NNs Learn to Program NNs (1991). Already existed and the idea as presented in the original paper doesn't work.
9. Learning Sequential Attention with NNs (1990). He uses the word attention but it's not the same mechanism as the one we use today which dates to 2010. This did not invent attention in any way.
10. Hierarchical Reinforcement Learning (1990). Their 1990 does not do hierarchical RL, their 1991 paper does something like it. This is at least contemporary with Learning to Select a Model in a Changing World by Mieczyslaw M.Kokar and Spiridon A.Reveliotis.
Please stop posting the ravings on a person who is trying to steal other people's work.