The difficulty of computing stable and accurate neural networks
pnas.org
pnas.org
Not surprisingly, the authors are mathematicians. They claim to prove that
* there are well-conditioned problems for which suitable DNNs exist, but no training algorithm can find arbitrarily good approximations of those suitable DNNs;
* it's possible to find approximations of those suitable DNNs only if we sacrifice digits of accuracy -- i.e., the approximations cannot be arbitrarily good; and
* there is a class of DNNs they propose, which they call "fast iterative restated networks" or FIRENETs, that solve undetermined systems of linear equations over the complex numbers, with a good blend of stability (robustness to adversarial samples) and accuracy (within the claimed theoretical limits).
Finally, the authors provide open-source code (A+ for doing that, but... Matlab!!??): https://www.github.com/Comp-Foundations-and-Barriers-of-AI/f...
Does anyone else here understand the work better than me? I would love an informal explanation that appeals to intuition.
Your statements appear to be a good summary of the paper.
Thanks. I found the authors' claims relatively easy to grok. What I'd like to understand, intuitively, is how they got there!
I skimmed through the proofs in the SI. Many of the tools they use are taken from the compressed sensing literature (robust null space property, sqrt LASSO paths, sample complexity for CS-MRI style problems...).
I can point you to additional reading, if you'd like.
The problem extends to any periodic functions. I am working on a blog post about it.
Traditionally, empirical solutions in the NN literature to address instability are regularization and dropout.
Also, adding layers seems to improve things. The famous example is XOR which cannot be trained by a single layer NN.
How do the theoretical limitations in the paper relate to these? if at all...
Innovation can come from anywhere. Either we can take time to evaluate something on its merits (if you have the competency and resources), or leave it be (if you don't), or we accept to be told what is true by institutions with their own interests, who will use their position to further leverage them. And in this path, knowledge is no longer free.
I will read further, but this paper is attacking a real issue with neural nets.
However, to me this appears to be serious work (whether or not it's remarkable). Crackpots often obscure their message, but here they state their premise quite clearly, and skimming through the arguments they seem sane and approachable.
I know this isn't the point you were going for, but this _is_ a heuristic that is deployed in some companies/institutions (at least the UK variant certainly is).
Which companies/institutions are you thinking of?
Granted, it could also be that this was authored by people outside of the field who.sont know the rules, but given how much research gets published each day, it's impossible not to rely on heuristics, especially since this paper was.posted without context.
Hinton's paper in 2006 on reducing dimensionality with NNs appeared in Science and nobody paid attention to that either at the time.