The paper's authors didn't succeed with Adam (which this article seems to have overcome) so I'm curious if they attempted this training method on any other datasets?
From the last paragraph of: https://openreview.net/forum?id=H1A5ztj3b
> Our experiments with Densenets and all our experiments with the Imagenet dataset for a wide variety of architectures failed to produce the super-convergence phenomenon. Here we will list some of the experiments we tried that failed to produce the super-convergence behavior. Super-convergence did not occur when training with the Imagenet dataset; that is, we ran Imagenet experiments with Resnets, ResNeXt, GoogleNet/Inception, VGG, AlexNet, and Densenet without success. Other architectures we tried with Cifar-10 that did not show super-convergence capability included ResNeXt, Densenet, and a bottleneck version of Resnet.
EDIT: I see now that they mention a few other datasets: Cars Stanford Dataset and Wikitext-2.