This is of no surprise, and well known. ViTs excel when data is plentiful, but perform less well than convolutions in lower data regimes. ConvNets on the other hand (resnets for example) don’t do as well in high data regimes. Imagenet is considered “low data regime” compared to what ViTs get normally trained on, most SOTA ViTs on imagenet are typically pre-trained on MUCH larger datasets (e.g. JWT300)