With that out of the way, here's my question:
Have you guys tried or managed to achieve learning super-convergence with attention or residual attention models?
I really hope that these results will help encourage more people to both see where and how super-convergence can be achieved. I suspect that we're only scratching the surface of what's possible.
I too hope these results encourage more people to see where/how super convergence can be achieved.
FWIW, trying to achieve it with fully attentional models is on my "R&D things to try" list at work.
Unrelated but it would be great if you could answer: In general, how important is momentum tuning and what are the heuristics for the same?
Unrelated but it would be great if you could answer: In general, how important is momentum tuning and what are the heuristics for the same?