Transformers in Medical Computer Vision
techblog.ezra.com
techblog.ezra.com
I didn't know about the HATNet mentioned in the article, but it looks an interesting paper to read.
There's a recent article[0] showing that a good training strategy with a ResNet can be better than novel architectures. From my experience, a ResNet-34 / ResNet-50 is usually enough (at least talking about classification), and the main bottleneck is your dataset, class imbalance, and handling out-of-distribution samples.