I'm going to write this out more clearly, because I'm still getting downvotes for my correct answer.
Why neural networks?
https://en.wikipedia.org/wiki/Universal_approximation_theore...
Can polynomials do this? (Yes)
https://en.wikipedia.org/wiki/Stone%E2%80%93Weierstrass_theo...
What is transformer and attention?
https://pathmind.com/wiki/attention-mechanism-memory-network
Attention = Polynomial (x2,x3 etc.)
Polynomial = interaction. VW flag -interaction
1 layer transformer = xx. (x^2)
2 layer tranformer = xxx. (x^3)
3 ... etc
What is reformer?
Transformer where LSH is applied.
One type of LSH is SimHash. ngrams of strings, followed by 32 bit hash.
Vowpal Wabbit -n flag for ngrams.
vw -interact xxx -n2 -n3 and you get ngrams + 32 bit hash doing SGD over a vector.
This vector is equivalent to a 2 layer reformer.
Non-linear activation is not needed because polynomials are already nonlinear.
So vw + interact + ngrams (almost)= reformer encoder. (if reformer uses SimHash, then they are identical).
Transformer/Reformer have an advantage, the encoder-decoder can learn from unlabeled data.
However, you can get similar results from unlabeled data using preprocessing such as introducing noise to the data, and then treating it as noise/non-noise binary classification. (it can even be thought of as reinforcement learning, with the 0-1 labels as the reward using vw's contextual bandits functionality. This can then do what GAN's do - climb from noise to perfection).