Findings from the Imagenet and CIFAR10 competitions
fast.ai
fast.ai
I'm not convinced that this is a real problem. The big tech companies have way more compute than anyone else, so they should do the large-scale experiments that no one else is able to do. It's true that you don't have to be particularly creative to come up with a lot of these experiments, but nevertheless they should be done and there is a lot of scientific value there.
I absolutely agree that more attention needs to be paid to smaller institutions since they're just as likely as anyone to come up with the next great idea. Double blind conference submissions help with this to an extent, but they're not a panacea. If you see an anonymous paper that has performed a thousand Imagenet experiments, is there really any doubt where it came from? And similarly, if they didn't perform a bunch of Imagenet experiments, is there really any doubt that it didn't come from one of the big players? So now you can mask a bias against small institutions as a subtler "you didn't perform enough Imagenet experiments" excuse. (Incidentally, this was the main reason that the paper by Smith & Topin was rejected from ICLR [1].)
EDIT: Full disclosure, I'm currently working on a set of not-so-creative (but still IMO important) experiments that use a large amount of compute at Google. So I have some bias here. :-)
With that out of the way, here's my question:
Have you guys tried or managed to achieve learning super-convergence with attention or residual attention models?
I really hope that these results will help encourage more people to both see where and how super-convergence can be achieved. I suspect that we're only scratching the surface of what's possible.
I too hope these results encourage more people to see where/how super convergence can be achieved.
FWIW, trying to achieve it with fully attentional models is on my "R&D things to try" list at work.
Unrelated but it would be great if you could answer: In general, how important is momentum tuning and what are the heuristics for the same?
Unrelated but it would be great if you could answer: In general, how important is momentum tuning and what are the heuristics for the same?