That's pretty interesting. It implies that the original accuracy rating might be legit. The concern is that we're chasing the imagenet benchmark as if it's the holy grail, when in fact it's a very narrow slice of what we normally care about as ML researchers. However, the fact that pretraining on JFT increases the accuracy means that the model is generalizing, which is very interesting; it implies that models might be "just that good now."
Or more succinctly, if the result was bogus, you'd expect JFT pretraining to have no effect whatsoever (or a negative effect). But it has a positive result.
The other thing worth mentioning is that AJMooch seems to have killed batch normalization dead, which is very strange to think about. BN has had a long reign of some ~4 years, but the drawbacks are significant: you have to maintain counters yourself, for example, which was quite annoying.
It always seemed like a neural net ought to be able to learn what BN forces you to keep track of. And AJMooch et al seem to prove this is true. I recommend giving evonorm-s a try; it worked perfectly for us the first time, with no loss in generality, and it's basically a copy-paste replacement.
(Our BigGAN-Deep model is so good that I doubt you can tell the difference vs the official model. It uses AJMooch's evonorm-s rather than batchnorm: [1] https://i.imgur.com/sfGVbuq.png [2] https://i.imgur.com/JMJ1Ll0.png and lol at the fake speedometer.)