Train TensorFlow models faster and at lower cost on Cloud TPU Pods
cloud.google.com
cloud.google.com
My understanding and experience is that it is not always trivial to get linear training speedups with additional machines.
It can be sometimes hard to get any speedup at all.
The difference with gradient descent can be described simply as:
* Single-machine training: take more steps
* Distributed training: take fewer but more confident and accurate steps. This could allow you to take bigger steps (learning rate) as well, but there is a limit to this as well.
They are not equivalent processes and it is an area of active research how to get an equivalent result with distributed training.
https://arxiv.org/abs/1811.03600
We looked at a bunch of different model architectures and datasets and found that you can get speedups from larger batch sizes up until a certain point, but that point is different for different datasets and architectures. The range of good hyperparameters is also narrower for larger batch sizes, which makes tuning harder.
I find that hard to believe. Can you be more specific?
They control the system attached to the chips. That includes the kernel and userland (glibc, libstdc++, etc.).
Very limited number of configurations (CPU, RAM, network, motherboard, any RDMA use) to qualify.
Monitoring, profiling, diagnostics, firmware, networking, security and automation follow the internal Google standards.
They can tune cooling to accommodate Google's motherboards and racks, whether it's air (v2) or liquid cooling (v3).
You can bet that they talk directly to the GCS backends through Stubby/gRPC, rather than sending HTTP traffic through the outside network and traversing GFEs.
Then there's all the other "mundane" stuff they don't need to worry about: packaging, manuals, warranties and user-facing RMA, multi-tenancy, etc.
Which, fair enough, is a pain.
Google's hardware layer is one thing; as btian mentioned, TPU also is deeply integrated with software infra e.g borg https://ai.google/research/pubs/pub43438
Besides, for special toys like TPU pods it's not clear that they would value adhering to OCP standards above anything else.
It probably takes a lot of effort to develop Windows driver, libraries, support for many kernel version etc.
That's easy to solve: "We only support Linux. Testing has only been done on kernel 4.13."
There is a huge amount of work that is put into ensuring that everything works reliably, with significant underlying infrastructure needed to achieve it. The problems you face at the level of scale Google works in on a day to day basis turns even the most mundane tasks into extraordinarily difficult algorithmic challenges that have to be solved.
So, by that reasoning, there is nothing that can exist outside of Google, because Google does things internally at scale.
Now think about how many different internal services like that a single complex ecosystem has to touch.
Saying "software is just a binary that needs to be made compatible with X platforms" is naive. It's like saying "facebook is just a bunch of UIs with form boxes, I could write facebook." Yeah, good luck.
If you look at an image of a Cloud TPU board [0], I'm not sure how it would even work with off the shelf servers. It's 100% custom for a Google Datacenter.
Nvidia also has the full supply chain set up, as well as things like driver support and all the things required to sell this kind of complex physical hardware.
Google also has a program to give research scientists free access to TPUs: https://www.tensorflow.org/tfrc/
[0] https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
(I work on GCP, but not on ML and TPU stuff)
Notice how Google did not submit benchmark results for several tests.
[0] https://blogs.nvidia.com/blog/2018/12/12/record-breaking-mlp...
So a Cloud TPU chip is slower than a V100 chip? Seems like that's the answer.
I’m also curious what the is the performance of the network interconnects in the TPU pod. I couldn’t find it documented, though I didn’t look super hard.
[1] https://github.com/uber/horovod/blob/master/docs/benchmarks....