Introducing the TensorFlow Research Cloud
research.googleblog.com
research.googleblog.com
If TPU teraflops are reported at fp16, then the number would be half.
Titan X offers 12 TFLOPs per GPU https://blogs.nvidia.com/blog/2017/04/06/titan-xp/
Where are you getting the extra 3 zeros from?
(Also 'up to 180 TFLOPs' is a bit misleading. As I read in some benchmarking posts by Google as well as by nVidida, a TPU is much faster than a GPU for doing inference, but they haven't released any data on training performance, the real bottleneck IMO).
Might also be TPU gen2 given the "64 GB of ultra-high-bandwidth memory" was not something IIRC the recent TPU paper was talking about.
And this is here now, while Volta is paper-launched and will be very scarce until late 2017/2018.
A GPU is also comprised of multiple chips (RAM, etc). I don't think "performance per discrete piece of silicon" is an interesting metric.
The closest comparison in terms of size would be 1 Volta DGX-1 (8x V100s) compared to 2 TPU modules (8x TPU2 chips).
Volta DGX-1: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
TPU Module: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
And for completeness, this is the size of a single V100: https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446...
You can see that 8x V100s are still more computationally dense than 8x TPU2 chips. Density is a very important factor in datacenter design.
http://hothardware.com/ContentImages/NewsItem/37034/content/...
(That's a P100 server, obviously, but it's the same kind of design. The P100 module looks a lot like the V100 module.)
The point of my previous comment is that it doesn't make sense to compare an entire TPU module (containing 4x TPUs) to a V100 board.
The closest comparison to a TPU module (with 4x TPUs) would be a dgx-1 board, which contains the nvlink bus that you mention, but also contains 8x V100 boards, hence why in my previous post I said you should compare the compute performance of a Volta DGX-1 (8x V100s) to 2 TPU modules (8x TPUs).
At the end of the day, it is simple, in a given area of space, you can get more compute performance from provisioning that area with V100s (in the form of using dgx-1s) versus provisioning it with TPUs (in the form of using TPU modules).
That's true if you're putting the chips in your own data center, because density affects TCO.
But assuming that Google does not sell TPU hardware, this isn't an important metric to anyone other than Google. The question that is important is, what does the shape of the curve plotting "dollars spent" versus "time spent waiting for my model to train" look like?
The shape of that curve is affected by TCO, and TCO is certainly affected by density. But there's a lot more to it than that.
But not to a single client.
Similarly, Google is not offering a 1000 TPU cluster to each client.
Your metric is strange to say the least.
The previous TPU was aimed at inference. The Cloud TPUs do training. But you're right - I haven't seen any publications about training performance yet. But they'll be available in cloud (currently in alpha), and once they go GA, I'd expect to see a raft of benchmarks by third parties. I can't wait. :)
V100 is 120 teraflops of tensor ops per chip: https://arstechnica.com/gadgets/2017/05/nvidia-tesla-v100-gp...
The closest comparison in terms of size would be 1 Volta DGX-1 (8x V100s) compared to 2 TPU modules (8x TPU2 chips).
Volta DGX-1: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
TPU Module: https://storage.googleapis.com/gweb-uniblog-publish-prod/ima...
And for completeness, this is the size of a single V100: https://cdn.arstechnica.net/wp-content/uploads/2017/05/34446...
You can see that 8x V100s are still more computationally dense than 8x TPU2 chips. Density is a very important factor in datacenter design.
Not everything has to be a data play in the big picture.
"Share the benefits of machine learning with the world"
It doesn't say who shares what :)
Not everything they do is about short-term data plays.
If the data can be used to improve the model it seems like data could be used to damage it, in theory.
Before anything else, the format and schema of the customer data would have to be analyzed and converted to data structures that match Google's internal models. While I imagine a computer could do it, I certainly would want a human to verify that the analysis makes sense.
Assuming this has been done, they are then at the mercy of the customer as far as whether the data is accurate, whether it is complete, how often it is updated, etc.
At the end of the day, I don't see how they would build a reliable business around arbitrary data structures which they have no control over. Information you can't trust is pretty useless.
Edit: They would also have to understand how the data was selected. Looking at a series of data points, you would wonder if all these are from Arizona, or all from the year 1976, or all from color blind individuals. Without understanding such limitations, making any sort of deduction from a dataset will just lead you the wrong way.
> how they would build a reliable business around arbitrary data structures which they have no control over
Google Search? The entire web could be described exactly like that.
†: Parses unreliably at best, Turing-complete at worst. (aka javascript if you didn't catch that)
Another point here is that Google Search isn't an authoritative source of information, it is up to the end user to inspect the returned links and decide if they can trust that site. This is something that I would not try to automate to the point that I could ask users for money in exchange, and if it can't be automated it doesn't seem like a great fit for Google.
Nah, the user's mood may be ruined if the search results are junk, but ruining the the model you're trying to build has vastly more costly consequences. I don't think the tech is there (just yet) to have some code simply ingest whatever comes its way, chew it up and use it well; required xkcd (today's!): https://xkcd.com/1838/
As far as I've heard, the hopes for Cloud TPUs in general are that people will think it's an amazing service worth paying for.
The hopes for TFRC are more complicated, as you might imagine when a company is giving away free stuff, but one of them, I'd put in the "long-term benefit altruism" category: It's quite hard to do some kinds of machine learning research in academia, because it can be fantastically expensive. We just blew $5k of google cloud credits in a week, and managed only 4 complete training runs of Inception / Imagenet. This was for one conference paper submission. Having a situation where academia can't do research that is relevant to Google (or Facebook, or Microsoft) is really bad from a long-term perspective, and one thing TFRC might do is help make sure that advances in deep learning continue at a rapid pace.
From watching the rate at which my deep learning colleagues get swallowed up by industry, I think it's a very valid concern, and I'm very supportive of all of the industry efforts we're seeing to try to address this. It's good for the entire research ecosystem.
(There are undoubtedly many other reasons, such as those noted below by minimaxir, but this is the one I personally feel the pain of.) Disclaimer: I get paid by Google part-time to do work related to this, but this is not any kind of official statement. The pain-of-ML-in-academia bit is purely from my hat as a CS professor.
To think that any entrant no matter the size could come into the space and successfully develop a device in such short order is amazing.
Not only that but have it used within their own infrastructure for some time, and later to allow outside external usage, again amazing from a planning, execution, product perspective. Ah shucks, from a hardware perspective too.
Especially coming from a company In the traditional sense which has no business being in this sector (excuses).
It seems like NVidia is an army, polished to pump "chips". And this Google special ops team kicked ass.
Is this a fair statement? I mean , this is Google. Hardware development is not easy or cheap, and they've been at it for years. They ponied up lots of money for AI starting as early as 2012, and this company is all about running huge jobs on lots of compute on giant data - it is hard to think of a company more suited to developing this kind of hardware (same for Facebook, in that sense).
Nvidia is the only current credible provider of DL hardware. If Google starts using TPUs then every other company will be forced to buy more Nvidia cards or they will be left behind.
Microsoft is the only exception here: they have some investment in a FPGA based ML solution. Even IBM uses NVidia as their DL solution on Power servers.
say "I'd like to introduce you to the concept of Intelligence Artificial."
delay 2
say "You might have heard of Artificial Intelligence, a term coined by John MiCarthy in the 1950s to label his study of difficult problems that his computer science lab was working on."
delay 2
say "Today we have a resurgence of Artificial Intelligence. You may or may not have noticed, but the Artifically Intelligent are now controlling your lives."
delay 2
say "From Siri, to Facebook, to the Googles, to the Amazon, your data is being fed into 'Machine Learning' algorithms at a rapid pace."
delay 2
say "Custom hardware is being designed this very minute to accelerate the training of these algorithms."
delay 2
say "That is what Artificial Intelligence is these days. Machines learning, about you."
delay 3
say "Intelligence Artificial turns this concept on it's head. By learning more about the machine, YOU, you can learn to control the machine."
delay 1
say "I"
say "A, not, A I."
delay 7
say "BusFactor1 Inc. 2017"
delay 1
say "Putting the ology into technology."
delay 3
say "Or, is it the other way around."
delay 4
say "I am the machine telling you about learning, this is my reference clip and this is where we are going"