I do keep up with literature and I do some applied research as well, so yeah, I see such things from time to time. The volume of papers is so intense though that unless there are other redeeming qualities if the paper does not use frameworks I already know (TF and PyTorch), I ignore it entirely. I wouldn't say I missed much that could help me in practice. One exception is Leslie Smith's work on cyclic learning rates and momentum modulation - he did it on some ridiculous setup, but it works for what I do.
I'm more surprised how many papers are written for tiny little datasets that you'd never use in practice, especially optimization papers. I mean, come on guys, I get it it's fast to train on CIFAR or fashion MNIST, but those results rarely translate to anything practical. And some papers are just plain not reproducible at all.
>> Google researchers are never going to use PyTorch en masse
IMO they should. It would easily double their productivity, and if Karpathy is to be believed their skin and eyesight would improve too: https://twitter.com/karpathy/status/868178954032513024?lang=...
>> I'd place my bets on Jax
As an ex-Googler, I'd place my bets in something else TBH. Google projects that aren't critical to Google's bottom line tend to deteriorate over time. Just look at TF. I'm not cruel enough to suggest it to my clients anymore, even though I could charge twice as much (because it would take twice as long to get the same result).