-Linear (transformer) complexity at training time
-Linear scaling with number of tokens
-Online learning(!!!)
The main point that made me cautiously optimistic:
-Empirical results on par with GPT-2
I think this is one of those ideas that needs to be tested with scaled up experiments sooner rather than later, but someone with budget needs to commit. Would love to see HuggingFace do a collab and throw a bit of $$$ at it with a hardware sponsor like Nvidia.