I Implemented Nyströmformer
github.com
github.com
I had to briefly look at the paper abstract, which explains that this is about solving the sequence limit of transformer text models:
>While beneficial, the quadratic complexity of self-attention on the input sequence length has limited its application to longer sequences -- a topic being actively studied in the community. To address this limitation, we propose Nyströmformer -- a model that exhibits favorable scalability as a function of sequence length.
That's cool — I'm looking forward to being able to process texts > 512 tokens in the future, and would be especially excited if that were possible for sentence-bert.