Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows
arxiv.org
arxiv.org
From the abstract: "Constructed entirely from standard ConvNet modules, ConvNeXts compete favorably with Transformers in terms of accuracy and scalability, achieving 87.8% ImageNet top-1 accuracy and outperforming Swin Transformers on COCO detection and ADE20K segmentation"
FAIR did great work with this paper.
FAIR definitely did great work with ConvNext, and I do hope to see more. There always needs to be people pushing unpopular paradigms.
[0] https://github.com/huggingface/pytorch-image-models
breakthroughs by definition upend conventional wisdom.
transformers currently represent conventional wisdom.
maybe transformers are indeed better than CNNs, but tech history is full of cases where primitive technologies take magical leaps when married to powerful hardware and intelligent system design.
low probability at this stage, especially given the impressive long-range dependencies and generalizability of transformers, but let's see what happens.
besides wightmann, who else do you recommend following for CNNs?
Hard to say tbh. Saining Xie was last on both convnext papers. I do know Trever Darrel always has interesting things going on in his lab. But I can't say anyone consistently though. I've done a few architecture papers but my focus is on generative modeling.
re your comment on chatgpt and citations, my sincere belief is that one of the more enduring benefits of LLMs, once researchers can eliminate/reduce hallucinations, will be to expose baseless claims and flimsy logic.
this won't eliminate all misinformation and misleading arguments, of course, but it will increase friction for bad actors and help good actors identify specious reasoning.
an automated BS detector if you will -- not unlike a noisy car alarm that deters theft.
what do you think?
if unified modeling isn't necessary, or even hurtful, a key advantage of transformers goes away.
For example, Swin loses some of the properties of self-attention but Neighborhood Attention doesn't[1] (these are both considered reductive attention types). Does this have a large effect? Might depend on your task.
Looking at non-classification tasks is important. This is especially important since ImageNet has a lot of issues with redundancy (e.g. there's a label "sunglass" (836) and "sunglasses" (837)), images with multiple labels that are valid, and more. I'd argue that once ImageNet accuracy is over 80% then the accuracy no longer strongly correlates with downstream tasks[2] (segmentation, detection) or even other tasks like generation. This is why it is really important for researchers to pay closer attention now and we can't just look at benchmarks. Doing so will hinder research. We could previously get away with this because classification previously strongly correlated with downstream performance and that the error rate in ImageNet was much larger than the improvements on accuracy.
Worse than that, some of the main benchmarks we use are highly effected by these biases. For example, convolutions learn texture and so using something like FID[3,4] can have plenty of issues that might not give an actual depiction of how good a network actually is at its task. This is even true for non-deep model based metrics[5] and so you have to be REALLY careful about how you evaluate things.
TLDR: be careful with evaluating benchmarks and evaluate things holistically.
===== Minimal Bib =====
[0] http://www.incompleteideas.net/IncIdeas/BitterLesson.html
[1] https://arxiv.org/abs/2204.07143
[2] If you're wondering how this can happen, it is memorization. E.g. if the target image has both rabbits, cars, and other things (real example) in it but the label is "car wheel" (479) then the network has to learn to ignore "car mirror" (475) and rabbits (330,331,332). But downstream tasks like detection and segmentation perform multiple classifications on a single image and using a over-fit backbone can hinder performance.
[3] https://arxiv.org/abs/2203.06026
[4] It is worth noting that FID is calculated form InceptionV3 weights, which was only trained on ImageNet-1k (full dataset is 22k) and had an accuracy <80% (top-5 < 95%). It is a bit weird this is used for datasets like FFHQ because there is no person label. Or even most LSUN classes because ImageNet likely has a texture bias itself (with animals and plants composing the majority of the dataset and texture being an important feature there). Evaluating models is fucking hard and a lot of this isn't internalized my many, especially outside the field.
The only reason convolutions are still used in modern (intelligently designed) ML systems is because it is not known how to build a sparse attention algorithm that achieves 2D and 3D locality and is also compatible with modern accelerators. Swin is an attempt at that, but it is something of a hack.
Both using a hierarchical transformer, adapting the transformer network architecture to vision tasks more efficiently.
Ah looks like people are still looking at them. https://bihy.medium.com/improving-convolutional-neural-netwo...