Advancing NLP with Efficient Projection-Based Model Architectures
ai.googleblog.com
ai.googleblog.com
Only on one task of text classification by 7 categories.
No expensive attention mechanism? This cuts down on number of parameters needed quite a bit, especially when sequences get very long.
I know what an RNN is but what is a QRNN?
Will need to read this more closely. I thought the key research direction now was improving the efficiency of the self attention block but maybe not.
Disclaimer, I am not affiliated to any of the authors