Sure!
TPUs aren't necessarily easier to use – it's about the same – but they're powerful. I've documented some benchmarks in this tweet chain, where I trained GPT-2 1.5B to play chess using a technique called swarm training: https://twitter.com/theshawwn/status/1214013710173425665
The power turned out to be from the fact that every TPU gives you 8 cores at your disposal. I never use the Estimator API. I just scope Tensorflow operations to specific TPU cores. Works great.
In terms of actual performance, I was delighted to discover that TPUs can be faster than GPUs when you use all 8 cores: https://twitter.com/theshawwn/status/1196593451174891520 (solution notebook: https://twitter.com/theshawwn/status/1205914446918492170)
It also gives you flexibility. TPUv2-8 can apparently allocate up to 300GB (!) if you don't scope any operations to any cores. Meaning, you run it in a mode where you only get 1 core of performance, but you get 300GB of flexibility. And then you can connect multiple TPUs together as described in the tweet chain, which quickly makes up the difference.
There is also the question of cost savings. A TPUv3-8 seems about as expensive as a V100. Which one is worth it? Well, it depends. In my experience a GPU is easier to use and quicker to set up if you only need one GPU of horsepower. But suppose you wanted to train a massive model in 24 hours. What's your best option? For us, it was TPUs.
The reason is subtle: It's hard to find any single VM that can talk to 140 GPUs simultaneously. But you can talk to 140 TPUs from a single VM no problem. And since you get 800MB/s to and from the VM, you can average the parameters across all TPUs very quickly.
This is similar to what TPU pods do internally. And while TPU pods are impressive, they are also impressively expensive. A TPUv3 pod will run you $192/hr at evaluation prices. Whereas you can play with a TPUv3-8 for $2.50/hr. You can also play with a TPUv2-8 for free using Colab: https://github.com/shawwn/colab-tricks
Yesterday I used that notebook to port forward Colab's free TPUv2-8 using ngrok, then trained using the new StyleGAN 2 codebase: https://twitter.com/theshawwn/status/1214245145664802817
I think a swarm of TPUs can cost significantly less than a cluster of V100s with less engineering effort.
That said, right now most codebases are designed to work with V100's. It will take time before TPUs widely proliferate. But speaking as someone who was once skeptical of TPUs and who has spent several months trying to discover their secrets, I feel that TPUs can get the job done quicker and easier than a GPU cluster. The hardware is also more accessible, since you can more easily spin up 100 TPUs than 100 V100s. But mainly I like that it's all coordinated from a single machine. It's conceptually simpler to debug and to implement.
If you run into any issues or have any trouble with TPUs, please feel free to ask here or DM me. I love talking about this stuff.
EDIT: In regards to usability, the new Jax library works with TPUs out of the box. Google seems to be heading in the direction of Jax. My initial reaction was "Not another library..." but first impressions were positive. It's not quite the React of ML – an idea which I hope to see soon – but it does seem easier for certain research purposes.
PyTorch also recently gained TPU support, and as far as I know they've put in some serious efforts to make sure things run quickly. As for how you use all 8 cores of a TPU using PyTorch, I haven't looked into it yet. But I'd be surprised if you couldn't. It seems unlikely that they would design an API that would hamstring you to just 1 out of 8 cores.