It also only applies to the continuous limit of non-stochastic GD, far from the real training methods used.
We don't gain any understanding either; understanding implies predictive power about some new situation, and I don't see any- and nor does the paper suggest them.
Looks like yet another attempt to attract attention by "understanding" NNs. Look, humans can't explain or understand how we drive, speak, translate, play chess, etc, so why should we expect to understand how models that do these work? Of course, we can understand the principles of the training process, and in fact we already do- the theory of SGD is well understood.