Transformers Are All You Need
pinecone.io
pinecone.io
LSTMs with attention are super interesting, and there are some neat self-growing hopfield networks with attention that could push into transformer style efficacy and efficiency gains.
Since then, this has been show to be untrue. Using more modern training techniques along with depthwise convolutions (https://arxiv.org/abs/2201.03545) results in equal if not better performance on vision tasks. Improved training methodologies have also been shown to boost the accuracy of ResNet50 - an 6-year-old pure convolutional architecture - on ImageNet-1k by over 5% (https://arxiv.org/abs/2110.00476).
Pure ViTs are also more difficult to train when compared with traditional convnets, although this has since then been somewhat remedied by Swin (https://arxiv.org/abs/2103.14030).
Although the number of parameters of a Transformer correlates with the amount of training data, this statement is misleasing. Precisely, the model wasn't trained on parameters but its parameters were trained on data.