Google announces a new generation for its TPU machine learning hardware
techcrunch.com
techcrunch.com
NVIDIA is now exerting pricing power to the point where they’ve decided to train their sales people to disregard the metric that is most important to customers: cost of training. Talk with one of their enterprise sales people and you’ll find they’ll say things like “FLOPS / $ doesn’t matter” to justify a 10x increase in price for their TESLA line. As history has shown, a monopolist sows the seeds of their own destruction and, by disregarding the metrics that matter, they alienate their customers.
Here are open source projects that you can contribute to to break the monopoly:
ROCm: https://rocm.github.io/
MIOpen: https://github.com/ROCmSoftwarePlatform/MIOpen
TensorFlow: https://github.com/tensorflow/tensorflow
I have done deep learning also and I don't see NVIDIA being an incumbent that is easy to be knocked out. I think we would need some distruption such as a new technique that is more performant or different design goals. Intel is still leading the way and I don't see NVIDIA being unseated with its head start anytime soon.
"Slowly grinds the mill of the gods, but it grinds fine."
Matrox didn't give much of a damn to 3D gaming, 3Dfx were bust doing all the wrong things. It was certainly not "easy", but don't seems hard or impossible either at the time.
https://github.com/plaidml/plaidml/tree/master/tile/hal
It's kind of cool to look anyway because it's a lot of value in less code than you might expect (complexity is in the layers above HAL).
If someone wanted to add support for Vulkan or something more ambitious like Qualcomm Hexagon or the RPi with QPU that would be pretty cool though.
How fantastic.
TF removes the silicon from the equation. So if Nvidia cheaper use but if TPUs cheaper use. If something new that is cheaper then move. TF is the key.
ArrayMancer: https://mratsim.github.io/Arraymancer/
Written in Nim, which looks and feels a lot like Python, but runs a lot like C++ or Rust, especially in the speed and FFI departments, and deploys like Go as a small single fast executable. Already supports CPU, CUDA and OpenCL. Still early, but with a lot of promise. Use is modeled on some mixture of PyTorch and Keras.
Are we sure about this? He specifically said that a "pod" would be 8 times faster than last year, not the TPU itself.
And the picture in the background showed what looked like to be 8 racks of 64 TPUs (or maybe 32?). Until now a "pod" was a single rack of 64 TPUs. So if the new definition for a "pod" is 8 times as many TPUs as it was before, the result is less impressive...
Is there any actual spec released?
How many chips sounds a lot more like a pissing contest. What if it was a giant chip with actually a bunch of chips inside? Who cares?
But, if you insist on going by the cost metric, you've already lost because you can buy nVidia GPUs and that's a lot cheaper than any of the TPU instances :)
Even more so with ML training. It is much more bumpy.
But you have a lot of other issues. Buying is going to hurt cash flow versus renting is a big one for small companies or startups not well funded.
Then also responsible for updating and stuck with old silicon. So for example we just got the TPU 3.0. So the cost will decrease.
But your cost of running what you buy is static. You are not be rewarded by improvements and they are happening quickly with the TPUs for example.
Just one year and we have the 3.0.
So, how much would 4k hours of 4x 1080Ti compute time cost me to rent? Or, to put another way, how many hours do I need to use to justify buy vs rent? Or, how many hours of 4x 1080Ti compute time can I rent for $6k?
But the other issue is you are stuck with the hardware. Google just put out the TPU 3.0 and you can use them without losing all your investment if you bought your own hardware.
The other is buying the hardware is much harder from a cash flow standpoint.
You provided some generic considerations, which don't apply to my situation, and I'm a fairly typical ML researcher. In fact, I know several startups doing DL research and they all have been buying hardware. Can you give me some examples where doing DL training in the cloud makes financial sense?
Also, you seem to imply that TPU is some kind of a miracle chip, even though we know very little about TPU2 and we know nothing about TPU3. It's actually pretty embarrassing that a general purpose V100 GPU is competitive with TPU2, which is an ASIC made for DL. If Nvidia ever decides to make a pure DL chip it will destroy anything Google can design. I mean, you can't be serious comparing Google to Nvidia when it comes to designing chips, right? That would be like believing that Nvidia could make a competitive search engine.
But the future is the cloud. I have fought people on this for well over a decade and thought everyone by now got it. Apparently not.
On TPU being a miracle chip not clear to me what a "miracle" chip would even be? Plus things keep improving and how could you have a static "miracle" chip?
What we do know is Google is doing WaveNet in production in real-time. We can hear the results. They are offering at a competitive price to traditional TTS.
We can also see the TPUs are about 1/2 the cost of using Nvidia.
Otherwise that is it. But I do not know your faith but personally I would not consider that a "miracle". But maybe just me.
Maybe Lewandowski would as I have heard he prays to a AI god or something like that.
The big difference is Google is working top down and Nvidia has to work bottom up.
So Google wants to roll out the best TTS there is to the world and at scale.
So they are going to from the top down to be able to achieve.
Clearly Google has a goal of creating the singularity and the Silicon is just a piece of the equation.
Versus I am not really sure what Nvidia goal is. Silicon is not for silicon sake. But rather a means to an end.
Plus Google having in production has the data to optimize in future versions where Nvidia just is not going to be able to.
I am long Nvidia and owned it for a bit. Was disappointed on how the stock traded after ER. Do think it will do fine as Google does all these incredible AI things. It is an alternative that will be good for them.
But for the very long term I would not hold. Google is a hold forever. Nvidia would keep on it.
BTW, would say Nvidia is now two generations behind.
Oh wow. Didn't realize you're one of those. I used to be a Kurzweil fan too, when I was 18. That phase usually passes by senior year in college. Well, I guess some people need longer to see reality. I recommend cutting back on pop tech news consumption.
Take it easy, and have a nice day!
BTW, I am old as in 50s. Have to explain to me in a little more detail what you mean by your post?
I am very curious but just not following?
Virtual machine pricing In order to connect to a TPU, you must provision a virtual machine (VM), which is billed separately. For details on pricing for VM instances, see Compute Engine pricing.
How much it cost you to complete that task is what matters. How it is done is here or there as long as get the precision.
We can see right now Google with their TPUs is about 1/2 the cost of using Nvidia with AWS.
https://blog.riseml.com/comparing-google-tpuv2-against-nvidi...
The 50% is a number you made up, and isn't based in any real benchmarks. For all other tasks that are not tensorflow tasks, the V100 is the only one that works.
But if we look at WaveNet we can see Google must be taking much larger margins with the TPU 2.0 versus Amazon using Nvidia.
Rolling out using a NN at 16k cycles a second and offering at a competitive price to the old way means Google TPUs have to be way more efficient than using anything from Nvidia.
It is hard to believe Google pulled it off. But if we look at WaveNet it suggests that TTS is a solved problem.
It will be how it is done for a very long time and just the NN will improve.
Nvidia honestly needs to get on their game. Google is running a 1000 mph. Iterating to the TPU 3.0 in just a year was a big surprise.
I suspect Capsule networks and dynamic routing which were invented by Google drove the TPU 3.0 but do not know.
Hope Google will share a paper now on the TPU 2.0 and their secrets.
Lucky for Nvidia they will share and Nvidia can copy. But just keeps Nvidia behind.
So WaveNet part of the Google Assistant and things like Duplex. But Nvidia silicon is just not going to make it possible. So Google does the TPUs as it just would not be possible to do at a reasonable price with Nvidia.
They are doing 16k cycles through a NN in real-time and competing against a far less compute intensive technique.
Now do not get me wrong I am long Nvidia and been for a while. Bit disapointed on the hit after an incredible earnings report.
I think they will do well from a investor marketing standpoint as really the only alternative to Google hardware.
But they have a fundamental disadvantage. They just do not have the applications like Google. They do NOT have the data to iterate like Google. It is why they appears to be 2 generations behind Google.
What is the goal of Nvidia? It is hardware for hardware sake? IMO, it should be driven top down and just do not see that happening at Nvidia. Is their goal the singularity?
Versus Google clearly wants to create the singularity and the silicon is in support. That is a very different calculus compared to Nvidia.
We can see it so strongly this week. The Google Keynote was all about AI applications. Then the TPUs are too support. Take a look at the duplex video for example. That is the focus and the silicon is what makes it possible.
But the more success for Google and presentations like Duplex this week and the buzz across the Internet helps Nvidia if the goal is investing. But from an actual technical solution it just points out the problem for Nvidia that much stronger.
So I care because I want to know if I will actually see this 8x speed improvement or if this is just marketing BS.
If NVIDIA tells me a 2080 is 8x time faster than a 1080, I know that a 2080 when released will be roughly the same price as the 1080 when it was released, and so I can expect to actually see a 8x cost/perf ratio improvement (if we put aside the fact that this are always best case scenarios)
If the 8x speed increase was achieve with 4x more TPUs (and roughly 4x the cost) like it seems it was, then this is just marketing bullshit, and we should really expect a 2x improvement in cost/perf ratio not 8x.
Will be interesting to see if the difference grows further with the TPU 3.0 or Google will just take larger margins.
Pretty sure you can get more than (or around) 8x speed up when comparing 8 V100 vs a single P100 for example.
- 64 TPUs per pod (11.5 petaflops)
- 4 chips per TPU (180 teraflops)
https://twitter.com/soumithchintala/status/99391656047246131...
We are doing data-center experiments - using oil cooling. It's suprising to see circuitboards dumped in ordinary oil and working away, no short circuits.
If you're interested, Alibaba has been doing some work along the lines of production immersion cooling.
What kind of oil are you using?
Mineral oil is ok for experimentation, but for long term material compatibility and fire risk I wouldn't recommend it.
FWIW I co-founded https://submer.com where we've developed an all-embedded computing immersion cooling solution that is virtually compatible with any kind of hardware (even fiber optics) and it's orders of magnitude more efficient than traditional data center cooling technologies.
Also - some oils while won't hurt your circuits, they might start dissolving the various plastic connectors on cables and board.
That will make a datacenter fire into something else entirely.
100F and below is considered flammable.
But what I wanted to know is performance in terms of joules compared to TPU 2s.
My hope is we get a paper on the TPU 2 now the 3 is released. We got the TPU 1 paper as they were releasing the TPU 2 I suspect.
Google doesn't need "developer buy in" for these to make sense -- they need better hardware for training and deploying deep learning models for their products, which is the overwhelming motivation. And, if they offer these to you, it's only because you're willing to pay for faster TensorFlow iteration.
Google has set the bar where others will have to match. Also looks like Google is now 2 generations ahead of Nvidia.
Would hope we get an even bigger price difference with this generation.
BTW, you are mixing up graphic generations and ML generations. Google never did or will graphics.
For ML Google is 2 generations ahead. Hard to imagine Nvidia ever catching up as so much of the AI breakthroughs are coming from Google.
A perfect example is WaveNet. Google is using a NN for audio in real time at 16k cycles a second. That is competing with using the old way of doing this. The Google approach gets a better result but people are only willing to pay too much more for a better result.
So Google had to price WaveNet competitively to the old approach which they have. I suspect that would have been impossible using Nvidia. The cost in running would have been just too high. The computation required for WaveNet is huge and the old way it is minimal.
Also we are going to need to really separate training with inference. They have different requirements and the costs in running are different.
Has Nvidia done an inference only solution? Or do they still always mix them?
"Google would have made a CPU to compete with Intel. "
I do a lot of surfing and have to say one of the more silly things to read. There is little gains today with CPUs. Why on earth would you invest in doing your own?
Google has ported all their stuff to Power so they have a backup and not reliant on Intel. But CPUs are NOT the future and would be crazy to invest into them at this point.
Processing is moving to TPU type processors and we will see more and more traditional things. A fantastic paper on this from Jeff Dean at Google I suggest you read.
https://arxiv.org/abs/1712.01208
One thing I love about Google is they do NOT reinvent the wheel. If there is something that can work like the Linux kernel they use. Versus having an ego of having to do yourself.
They did all their own network silicon by hiring the Lanai team several years ago because there was nothing that could work for their needs. They then created a determinate network stack so they could do Spanner. There was nothing off the shelve. Heck they laid their own fiber under the ocean to make possible.
Spanner is the first and only horizontally scalable RDBMS solution to exist. Google beat the speed of light by using their custom silicon, custom stack that is determinate, their own fiber. Then using atomic clocks with GPS to remove the latency from speed of light out of the equation. Highly recommend the papers including the Spanner paper.
https://ai.google/research/pubs/pub39966
The network uses custom silicon inside and then they put the intelligence on out edge. Then they control exactly the traffic so they get a determinate result. This also allows them to use far cheaper hardware as do not need to over provision as they never drop packets on the ground or need much memory for buffers.
Google would have used Nvidia in a second if they could do what they need.
But what is being done in AI (ML) today requires unique silicon. Just not possible without as the power footprint would be prohibitive.
I do expect Google more and more to leverage on client ML where needed with their PVC. So it is NOT just the TPUs but for somethings you are going to need on client.
Take the new voices as an example. They are doing 6 voices as each has a different model. The cost in switching models is prohibitive in the cloud.
So you will see them eventually move to some things being done on the client. But they are so far ahead they have tons of time.
I would not be surprised if we do not seen anyone else do Wavenet for a couple of years. What Nvidia should be working on is finding a way to do WaveNet at a reasonable cost for Amazon and others.
But the long term is the big cloud providers will do their own silicon. Google set the path with the TPUs and now the others will have to do the same to keep up. But just makes sense as the cloud providers have the data to iterate.
MS is using a FPGA which was foolish, IMO. Never going to get there with that route.
BTW, you need to realize how this works for inference. You keep the model in memory.
Google invented Capsule networks that use dynamic routing and are going to cause different patterns in memory access and would be curious if the TPU 3.0 were optimized for such?
They have a huge advantage of inventing the algorithms and then get to make their own silicon optimized.
Something Nvidia really needs to be able to better compete, IMO.
BTW2, Google using Power did include the excellent Nvidia memory interconnect as part of the spec.
Not sure did Google license or does Nvidia give away?
The spec is called Open Power if not aware?
https://www.theregister.co.uk/2016/04/07/open_power_summit_p...
Plus could they do the new text to speech at a competitive price using Nvidia Silicon?
Google approach gets you a better result but the computational requirements would be huge compared to the traditional approach. You have to get the compute cost down as far as you can.
No conspiracy here, we were simply late to market on the V100 Beta. As many have noted, we also hold a really high reliability bar before going to Beta (“Google’s Beta is like AWS’s GA” though I wouldn’t claim that in all cases) and we’ve be in Alpha for months. If you’d like to be included in Alpha offerings drop us a line.
As I’ve said in previous threads: Compute Engine intends to be the best place for computing. That includes the kind of workloads TPUs are tuned for, but it also includes all the stuff that GPUs excel at. That’s not going away, and GCE isn’t favoring one over the other. We sell infrastructure, and we don’t try to (overtly) pick favorites, we let our customers do that.
Wow, that's so ... I don't even have a word for it.
It is like how they ported to Power for their stuff if needed.
Look at the shortage of late with Nvidia. Google does not have to deal with it.
But also I suspect they could not do WaveNet without their own silicon as the cost would be prohibitive. Using a NN for speech and competing against traditional approaches and have a competitive price is going to require really efficient silicon.
GPU shortage equals four-month wait time for buyers"
https://www.theregister.co.uk/2018/03/30/nvidia_geforce_chip...
But you are missing my point. Google is not dependent on Nvidia.
Really does not look like Nvidia could do what they need anyway. Doing their new text to speech on Nvidia I would not expect to have a price point to make viable.
Please show me an article where there's a shortage of Tesla cards.
But really now it is just Nvidia does not have the ability to meet Google's needs.
I am not aware of anything from Nvidia that could do 16k cycles through a NN in real-time at a cost that you could roll out at scale.
Google priced WaveNet competitively to the old way of doing. You can only do that with a very low cost for computation and the power required which Nvidia could not match with the TPU 1.0 let alone now they have the 3.0.
But right now they have what they need.
BTW, I would be curious on the TPU 3 design considerations to better support dynamic routing with Capsule networks. I suspect creates different memory access patterns and where the big power savings are at today.
They shipped Deep Learning Accelerating Card SC1 at $589. You can't buy it right now because it is sold out.
We're long overdue for a general-purpose CPU with say 1024 cores, that avoids a central main memory, where each core can be independently programmed just like any other CPU. Google's may count as a somewhat general-purpose DSP, which is definitely a step forward. But no matter how mature or mainstream a framework like TensorFlow gets, it can never replace full programmability.
Without seeing the internals, I'm going to have to give this a nay vote for now. There are many other rather exciting problems that need to be opened up to a new generation of tinkerers. Off that top of my head, it's things like: content-addressable memory to provide high data locality (evolving various interconnects instead of hardwiring them), exploring other types of general vector processing like the kind MATLAB/Octave uses, and exploring other hill-climbing algorithms than backpropagation/neural nets.
I picture something more like network topology-agnostic Docker containers programmed in Elixer/Erlang/Go that can act as semi-autonomous agents and switch into various modes in order to solve the problem at hand. I just find that a much simpler metaphor to work with than OpenCL/CUDA/TensorFlow. Yes it would take more silicon and would probably violate YAGNI, but only full programmability gives us the freedom to explore the problem space at the level that's going to be required to implement artificial general intelligence.