I'd recommend doing a performance-per-dollar comparison before drawing this conclusion.
It’s hard to compare perf per dollar. Your electricity costs and GPU costs maybe vastly different from mine.
Now the real problem is, can you actually get a TPU pod in practice?
48 to 1 hour would have allowed us to run a few experiment per work day. That would be huge.
If you told me that the next generation of TPUs could run the current 48 hour GPU training set in 1 hour, that would be something.
But 2x faster won’t make a difference. Generally, startups for a new chip need to demonstrate 100x improvement over the general purpose approach to get funding. So this seems like another Google vanity project.
If you really wanted to get from 48 hours to multiple experiments per work day, you could (probably) do that right now by scaling horizontally. However, the cost becomes a major issue when you do that on GPUs. This is where the TPU shines since you can scale out without cost crippling you.
I work on training on GPUs and the TPU is definitely an incredibly useful piece of hardware, not a vanity project.
But, TPUs are not standard, can’t be used for any other usecase. So those savings might not actually be realizable when everything is taken into account.
TPUs do have the memory capacity advantage though (over 2080Ti).