Notably, this result would be very hard for Microsoft to achieve, or indeed even reproduce fully in-house, because it requires more memory than GPUs have. GitHub mentions that TPUs are pretty much table stakes to train this.
https://medium.com/tensorflow/fitting-larger-networks-into-m... https://medium.com/huggingface/training-larger-batches-pract...
could you elaborate? TPU V3 unit has 16GB of memory, and old V100 also has 16GB of memory. Plus TPU has extra memory consumption for mandatory tensor padding, which GPU doesn't have.