It seems since submitting that they are no longer the leader on the GLUE leaderboard - https://gluebenchmark.com/leaderboard/
Microsoft: We beat human performance and lapped everyone else!
XLNet: Hold my beer.
Code: https://github.com/zihangdai/xlnet
Description: https://towardsdatascience.com/what-is-xlnet-and-why-it-outp...
https://medium.com/tensorflow/fitting-larger-networks-into-m... https://medium.com/huggingface/training-larger-batches-pract...
could you elaborate? TPU V3 unit has 16GB of memory, and old V100 also has 16GB of memory. Plus TPU has extra memory consumption for mandatory tensor padding, which GPU doesn't have.