Luckily they’ll fix the issue within 24h, which is when TPUs preempt. :)
Luckily they’ll fix the issue within 24h, which is when TPUs preempt. :)
I manage an enterprise machine learning team and we do tons of stuff on GCP. I’m dead serious: there is not a single use case where it makes sense to choose TPUs unless the main value you’re seeking is just the entertainment factor of using Google’s new thing.
Since your post didn't actually include any details, I did a search and immediately found an article[1] where TPUs worked better for their particular use-case. I suspect I could find many such reports (and probably some opposite reports too).
It's unfortunate they didn't work for you, perhaps you should give them another shot with a different model. I'd recommend using Cloud's examples as a starting point.
[1] https://medium.com/bigdatarepublic/cost-comparison-of-deep-l...
When you use them right, a TPUv3-8 gets equivalent perf to a cluster of 8 V100s.
I was astounded. I trained StyleGAN 2 from scratch at 1024x1024 in 2.5 days. nvidia took 7 days for their official model. Granted, I used a v3-32, not a v3-8, but performance seems pretty similar.