> ViT models are outperforming CNNs in terms of computational efficiency and accuracy, achieving highly competitive performance in tasks like image classification, object detection, and semantic image segmentation.
Since then, this has been show to be untrue. Using more modern training techniques along with depthwise convolutions (https://arxiv.org/abs/2201.03545) results in equal if not better performance on vision tasks. Improved training methodologies have also been shown to boost the accuracy of ResNet50 - an 6-year-old pure convolutional architecture - on ImageNet-1k by over 5% (https://arxiv.org/abs/2110.00476).
Pure ViTs are also more difficult to train when compared with traditional convnets, although this has since then been somewhat remedied by Swin (https://arxiv.org/abs/2103.14030).