If the claim holds, the implication is that layers can quickly learn much of what they need to learn locally, that is, without requiring backpropagation of gradients from potentially very distant layers. I can't help but wonder if this opens the door for more efficient asynchronous/parallel/distributed training of layers, potentially leading to models that update themselves continuously (i.e., "online" instead of in a batch process).
I wouldn't be surprised if the claim holds. There is mounting evidence that standard end-to-end backpropagation is a rather inefficient learning mechanism. For example, we now know that deep neural nets can be trained with approximate gradients obtained by shifting bits to get the sign and order of magnitude of the gradient roughly right.[1] In some cases it's even possible to restrict learning to use binary weights.[2] More recently, we have learned that it's possible to use "helper" linear models during training to predict what the gradients will be for each layer, in-between true-gradient updates, allowing layers to update their parameters locally during backpropagation.[3] Finally, don't forget that in the late 2000's, AI researchers were doing a lot of interesting work with unsupervised layer-wise training (e.g., DBNs composed of RBMs, stacked autoencoders).[4]
This is a fascinating area of research with potentially huge payoffs. For example, it would be really neat if we find there's a "general" algorithm via which layers can learn locally from inputs continuously ("online"), allowing us to combine layers into deep neural nets for specific tasks as needed.
[1] https://arxiv.org/abs/1510.03009
[2] https://arxiv.org/abs/1602.02830
[3] https://deepmind.com/blog#decoupled-neural-interfaces-using-...
[4] https://www.iro.umontreal.ca/~lisa/pointeurs/TR1312.pdf
EDITS: Expanded the original comment so it conveys better what I actually meant to write, while keeping language as casual and informal as possible. Also, I softened the tone of my more speculative observations.