In fact the article was published on June 20, while the XLNet submission that dethroned them was already on June 19. I guess their publishing pipeline doesn't allow last-minute amendments.
Microsoft: We beat human performance and lapped everyone else!
XLNet: Hold my beer.
Code: https://github.com/zihangdai/xlnet
Description: https://towardsdatascience.com/what-is-xlnet-and-why-it-outp...
https://medium.com/tensorflow/fitting-larger-networks-into-m... https://medium.com/huggingface/training-larger-batches-pract...
could you elaborate? TPU V3 unit has 16GB of memory, and old V100 also has 16GB of memory. Plus TPU has extra memory consumption for mandatory tensor padding, which GPU doesn't have.