Happy to answer questions!
Happy to answer questions!
EDIT: you don't get sustained use discounts, either, at the moment. You can get either for GCP GPUs, though. Perhaps that will change once TPUs are out of beta?
Any idea of how much variation in accuracy you get on different training runs of the same model on the same hardware? My understanding is that model quality can and does vary from one run to the next on these kinds of large datasets - from a single observation, it's hard to know if the difference is real or noise.
Tracking down bugs in convergence is really costly in these settings. We had a problem in pre-processing that took us quite a while to figure out...
But what's going on when some of the implementations of a standard algorithm don't converge, and different hardware has different accuracy rates on the same algorithm? Are DNNs really that flaky? And does it really make sense to be doing performance comparisons when the accuracy performance doesn't match?
Is the root problem that ResNet-50 works best with a smaller batch size?
And how do you do meaningful research into new DNNs if there's always an "Maybe if I ran it again over there I'd get better results" factor?
Thank you.
Note also, that the ~2% performance difference is only on one model (ResNet-50) and cannot be generalized to all workloads/all of deep learning (at least not without further proof).
That is not all that close is it?
the TPU implementation applies very compute-intensive image pre-processing steps and actually sacrifices raw throughput
Thanks
In terms of how much compute power the TPU pre-processing needs I only have very rough numbers: I ran the same pre-processing while training ResNet-50 on a node with 4 GPUs and it was consistently utilizing >22 CPU cores (including all of the other CPU-tasks while training).