The Reformer – Pushing the limits of language modeling
colab.research.google.com
colab.research.google.com
I wonder how much of a limitation it poses that the gradient requires some massaging to preserve efficient training; does that only work for some kernels, or can it be automated for arbitrary kernels?
[edit] Linformer (https://arxiv.org/pdf/2006.04768.pdf) is a different project from the one linked in https://linear-transformers.com/ (Transformers are RNNs https://arxiv.org/pdf/2006.16236.pdf).
For me the greatest trick is improved runtime compared to older seq/RNN techniques
In particular at least the Roberta model by Facebook is already improving significantly upon XLNet.