Under the Hood of Google’s TPU2 Machine Learning Clusters
nextplatform.com
nextplatform.com
Regarding the power consumption of the TPU2 racks and whether or not Google will continue to use them: as long as they're more power efficient than GPUs of course they will continue to be used. That is most likely the only reason the TPU exists in the first place. Running the same workload on GPUs would consume a few Mega Watts if the TPUs consume 500 KW. And Google has lots of workloads like that so the savings must be enormous.
But GPUs are a moving target and even if for now the TPU seems to have the edge on power efficiency that could easily change.
Another thing that does not make much sense to me is that if you were to use systolic array processors with a certain limit that it would make very good sense to figure out a way to daisy chain them so that multiple units can be combined into larger arrays.
That may very well be the function of that chip in the center, some kind of crossbar to allow for easy linkage of single TPUs into larger fabrics.
If I look at a GTX 1080 Ti and optimistically assume that I can keep it completely busy 24/7, operating at the 250 watt TDP, it would take 4.5 years for the retail electricity cost to match the initial purchase cost of $699. (I pay a bit under 7 cents/kWh). Now I do live in one of the cheapest-electricity regions of the US, but I would also assume that big data center operators are building facilities where electricity is cheap. And big customers get lower rates than households. I wouldn't be surprised if a GPU cluster's hardware is considered obsolete before its electricity costs match its initial hardware costs.
If Nvidia is smart, they consider only marginal, not fixed costs when pricing their products.
If e.g. Q(p) = 1/p^2 (price elasticity of demand of -2), then profit is 1/p - c/p^2, which is maximized at p = 2c.
Note how optimal price depends on cost.
Often bundling decreases cost. Bundling also has the additional advantage of producing more demand when selling excess capacity a la AWS.
Google wants a die as efficient as possible All of those take up die space, creating a larger processor. That increases the cost of defects, as well as overall manufacturing costs.
For Amazon, you may be right, but for Google, they only want a specific function.
Failing math, let's hope AMD'em Vega GPUs don't completely suck.
Despite my general server / hardware comprehension being on the low end of intermediate this article was extremely digestible and helped close some gaps for me about Google's strategy and how/where to think about deploying cloud-ML in the future - love seeing content like this on HN.
The article is a great piece of analysis.
Could this instead imply Google is working on its own library of low-level math primitives just for Tensorflow, and if so, how long would that take before it's competitive performance-wise? At any rate, having to support a different computing platform / API would be another blocker to general adoption.
This undertaking isn't something to be taken lightly; CuBLAS has some pretty cutting edge architecture-specific optimizations for batching operations for matrix multiplication that came from several years of research - and is arguably a massive competitive advantage of NVIDIA over AMD. Depending on the development state of such an API internally within Google, it could mean that the Cloud TPU isn't going to be ready for wide-spread commercial use for a good while, and is very much still in the research-and-development phase (which could explain why they're only opening it up for the research community right now).
And if you caught any of their talks on [Tensorflow at Google IO](https://www.youtube.com/watch?v=5DknTFbcGVM), it seems like their goal is to provide ultra-high level APIs like POST/GET, and high level python ones like Tensorflow pre-built and trained models. At that point Google can control the engineering, and as long as you run on their cloud platform you just have to worry about the high-level. I'm definitely not a fan of vendor lock-in, but it seems like a interesting product, and I'm curious to see what Google does in this space.
I really worry about all these reduced floating-point representations when they are made use of by people who mostly understand deep learning through tinkering with existing tensorflow tutorials.
FP32 seems like a relative sweet spot with sufficient dynamic range to let most amateurs avoid getting trapped in the weeds. I could probably be persuaded to believe that FP24 is sufficient as well.
But I suspect that once you get down to 16 bits or so you have to do all sorts of stochastic rounding / dithering that is beyond the skill set of most data scientists. And that's because I suspect many of them don't really get dynamic range.
Which then leads me to believe that what we really need here is the equivalent of R for machine learning.
https://www.nextplatform.com/2017/05/17/first-depth-look-goo...
In portrait mode, the text scrolls off the side of the screen so I have to continually pan back and forth to read an article. In landscape mode, a very thin column of text is used. It's as if the landscape res media query is being applied to portrait view and vice versa.
Edit - oh hmm looks like manually zooming out a bit in portrait will fit the whole article. Was sure that didn't used to be the case in earlier days. Maybe it got fixed!