Open-Source Simple-ViT Implementation
github.com
github.com
An update from some of the same authors of the original paper proposes simplifications to ViT that allows it to train faster and better.
Among these simplifications include 2d sinusoidal positional embedding, global average pooling (no CLS token), no dropout, batch sizes of 1024 rather than 4096, and use of RandAugment and MixUp augmentations. They also show that a simple linear at the end is not significantly worse than the original MLP head.
Simple ViT Research Paper: https://arxiv.org/abs/2205.01580
Official Github repository: https://github.com/google-research/big_vision
Developer updates can be found on: https://twitter.com/EnricoShippole
In collaboration with Dr. Phil 'Lucid' Wang: https://github.com/lucidrains