A stack of feed-forward layers does surprisingly well on ImageNet
arxiv.org
arxiv.org
This was a short write-up of a set of experiments exploring the importance of attention in transformers. I'm glad to see that people are still reading it and hopefully finding it interesting.
Side note: the title should probably have (2021) in it, as this was posted to Arxiv in May 2021 (concurrently with a number of similar works such as MLP-Mixer).
I'm not sure how to prevent this, maybe you can preregister your programmers?
I hadn’t seen that result before, definitely interested in related work
The CIFAR experiments I mentioned were https://arxiv.org/pdf/1806.00451.pdf. It doesn't contain this argument (unfortunate wording) but appears to support it well.
Looks like the attention layers aren't as important as everyone thought because using a similarly sized feed-forward layer works almost as well (74% vs 77% top1 accuracy)
If a network is the same size, less computationally demanding, and gives you 3% improvement, it seems extremely worthwhile, especially given the (mostly) diminishing returns of just adding more weights/layers.
Whether we need attention or not is a more interesting question on seq2seq models on text data.
https://github.com/fawazsammani/awesome-mlp-mixer
"MLP Mixer: An all-MLP Architecture for Vision" is #1 on the list. The paper we are commenting on is #2 on the list. They were uploaded to arxiv two days apart!