GPUs for Google Cloud Platform
cloudplatform.googleblog.com
cloudplatform.googleblog.com
Shameless plug, if you want raw access to a GPU in the cloud today, shoot me an email at daniel at paperspace.com We have people doing everything from image analysis to genomics to a whole lot of ML/AI.
This looks like it might be a cool solution to that problem. Use iOS most of the time (I'm a writer) and the Paperspace VM on the iPad for coding (I'm also a wannabe developer).
Only fly in the ointment is that it's not available in Europe, and I'm worried about latency over a 4g connection.
For coding, you don't even need a GPU. Just use RDP. Big companies use this all the time (thin clients + central server farm, also called VDI).
If you want Linux remoting, I recommend either x11vnc or xrdp.
It says in big font "Starting at $5/month". Then I clicked "Pricing" and the cheapest I see is "$15/month". What's the $5 plan?
There was 0 lag when connected directly over the ethernet. It was seamless.
Though I should add that the paperspace server I was connecting to was located less than 50 miles from my laptop.
This better be worth it! :P
Wow, that's impressive. One thing I've loved about GCE has been the custom sizing. This takes it even further, so we don't have to buy what we don't need.
Looking forward to seeing the pricing on this. Looks like they're going to heavily compete with AWS on this stuff.
Disclosure: I work on Google Cloud (and want to sell you some GPUs!)
Compare this was something like oracle, which simply can't consume the unused resources in order to discount the hardware effectively. They can't beat GCE/AWS at the cloud game until this changes.
Actually not sure about Google either – their preemtible instances are sold at an 80% discount, so that puts an upper bound on the usefulness of these machines for them.
http://www.networkworld.com/article/2891297/cloud-computing/...
Here's how I think AWS and GCE work...
Imagine you buy some hardware. You have n=1 users, and so the variance on your compute needs is very high. You buy for peak use, so your median utilisation is low.
Now imagine you have n=10,000,000 users. Again, you stock for peak use...But the variance is now much much lower. Your median utilisation is now much better.
Let's quantify your workload as "resource-seconds", what many of Google's services (and some AWS services, like Lambda) allow you to do is maximize resources and minimize seconds (assuming efficient parallelization of workload, which is afforded by Google's super-awesome Networking). Cost is the same to you, but you get your results faster.
This is afforded by the scale and by the multi-tenant nature of underlying services. This model can only work when there's lots of folks using such a service.
(work on Google Cloud Big Data)
Curious because if I pay for using a P100 for 2 hours - it is 100% available to me and not virtualized of any sort right?
However, if I use it only for 2 hours that day, and no one else were ready and waiting to buy the time on those, then Google uses it internally until...
> GPUs are offered in passthrough mode to provide bare metal performance. Up to 8 GPU dies can be attached per VM instance including custom machine types.
When you're using it, it's 100% yours.
Disclosure: I work on Google Cloud (and pitched in on this!)
What "unused resources" are used for "their own compitation"?
How does this "pay for itself"?
I've never worked at google. Do they just let each team go with their own crazy hardware setup? Or maybe two separate systems, one internal and one external? That seems a little wasteful (they have the money, so whatever).
But wouldn't programming against a big unified api be a bit more, well, sane? I can see really sensitive stuff locked away on its own hardware, but the search front end? why not serve some JS or provide chrome downloads from a giant pool of common hardware?
Things grow organically, and they might be technically locked into a specific layout right now, that's totally understandable. But, unused capacity is just capital depreciating steadily away with no gain. Soaking up every spare cycle doing something useful should probably be somewhere in the company goals.
You can get an answer on your laptop in an hour. You'll probably get a better answer throwing 1000 hours at your algorithm though, and a better answer throwing 10000 hours.
You can do a lot of neat things if p=np. Since no one knows if that's true, we're stuck with bad big O. More computer time is the only way out right now.
The point is, I'd be disappointed in google researchers if they didn't have a huge amount of unscheduled workload.
"But Witold Jarnicki and David Konerding already did that: they wrote a C++ program that built a table of Cz(modp)Cz(modp) up to 5000500050005000 , and, in parallel across thousands of machines, searched for A,BA,B up to 200,000 and x,yx,y up to 5,000, but found no counterexamples. On a smaller scale, Edwin P. Berlin Jr. searched all CzCz up to 10171017 and also found nothing. So I don't think it is worthwhile to continue on that path."
I wonder what kind of device driver does GCE use with AMD, the new ROCm?
What about Power8 + NVLink harware? Does anybody know if the current NVIDIA GPUs, in particular the P100s are all on x86?
We will be working with both vendors to make sure we highlight which drivers work most reliably on our stack though (virtualization + GPUs is way too rare!). We want to work closely with both vendors to do the qualification, so you at least know what is known good.
Disclosure: I work on Google Cloud.
> If those bytes happen to be Debian 8 with some AMD driver that supports the 9300, awesome ;).
happens to support might alone not be enough given how much difficulties AMD has been having. I just hope that there will be enough customer demand for their SPFlop cruncher GPUs for interested parties -- including Google -- to pitch and contribute to their ROCm stack.
http://hothardware.com/news/amd-rocm-13-adds-support-for-pol...
They are revamping their entire linux stack for HPC, heterogenous compute.
"New Linux Driver and Runtime Stack optimized for HPC & Ultra-scale class computing" sounds great, let's hope this time they can do it.
I'm still very worried that AMD will fail, not because of the hardware, but due to their lacking software stack.
"GPU-Accelerated Microsoft Cognitive Toolkit Now Available in the Cloud on Microsoft Azure and On-Premises with NVIDIA DGX-1"
DGX-1 is powered by 8 x P100
[1]http://nvidianews.nvidia.com/news/nvidia-and-microsoft-accel...
https://www.nimbix.net/blog/2016/10/04/ibm-nvidia-powerful-g...
Nice to have a full cloud provider with a lengthy feature set AND P100s though.
> GPUs are offered in passthrough mode to provide bare metal performance. Up to 8 GPU dies can be attached per VM instance including custom machine types.
So when your VM boots up, the GPUs you attach will not be shared with others. I agree that the state of GPU virtualization is not yet where it needs to be to make me personally comfortable enough.
Disclosure: I work on Google Cloud (and worked on this a little).
For models you write yourself and have the Cloud ML service train, they will use our regular VMs today and as this blog post says, GCE will soon have GPUs (and therefore, Cloud ML will expose them, too).
Disclsoure: I work on Google Cloud (and sorry for the confusion!)
https://aws.amazon.com/about-aws/whats-new/2016/09/introduci...
Google: up to 8 K80s and ~200 GB of RAM but much more flexible
At the same time, in many other complex applications the performance/W/$ advantage has not been all that clear, especially since GPUs have been quite behind in process technology. A good example is a simulation code I work on where a 2-4x speedup from GPUs, which (due to inherent overhead of using accelerators) gradually vanishes in strong scaling, means that when price is considered, only cheap consumer cards can imrpove the performance/buck metric significantly (that's time-to-solution/buck in our case), professional cards don't offer significant significant improvements other than performance density [3, figures 6,7].
There are plenty more examples that NVIDIA collects [4], but always take the claims with a grain of salt, especially the ones with "incredible" >10x speedup claims :)
Last, I'd note that with the recent 14-16nm jump, the gap between the traditional CPUs and the simpler accelerator architectures has increased (I've just seen MAGMA BLAS on Tesla P100 results which show >10x GFlops/W, more than double that of the previous arch) and I expect it to keep increasing partly due to the manufacturing/process technology gap shrinking between Intel and the rest and partly due to architecture and programmability improvements of accelerators.
[1] https://www.nvidia.com/object/gpu-accelerated-applications-t... [2] http://on-demand.gputechconf.com/gtc/2015/presentation/S5476... [3] https://www.academia.edu/13753737/Best_bang_for_your_buck_GP... [4] https://www.nvidia.com/object/gpu-applications.html
Machine Learning is one such field, batch-rendering for CG studios etc is another.
You could throw 100 CPUs at something that perhaps 1 GPU might do in the same time (made up numbers but it gives you an idea of why people want it)
What is your concern with virtualization in this context?
Disclosure: I work on Google Cloud (and for a long time specifically Compute Engine).
I think you're alluding to security issues with renting bare metal hardware to potentially untrusted customers, right? I've often wondered about the security implications of that - think backdooring BIOS/UEFI or other component firmware. But there are many large providers that rent dedicated hardware without a hypervisor. Are they just ignoring these concerns, hoping that this kind of tampering would be rare anyway?
> What is your concern with virtualization in this context?
I'm not the parent poster, but performance was probably the concern. There's still some (marginal, these days) overhead with virtualization in general, and VT-d specifically.
In this context though of GPU compute, you are just talking sporadically to a PCIe attached device directly. There isn't any virtualization or sharing of the device. And yes, I assumed that the complaint would be performance, but especially with just a pass through GPU, it's going to be honestly funny to find the right "oh wow this is 20% slower because virtualization specifically".
That's what I figured. I still wonder what other large providers (SoftLayer, OVH, etc.) do about this. Unfortunately, I suspect the answer is "nothing, we just hope/assume nothing will happen."
> In this context though of GPU compute, you are just talking sporadically to a PCIe attached device directly.
I know; I'm familiar with passthrough. It's not completely "free," but the overhead is very minor. I don't think anyone will be able to complain with a straight face about its performance. :P
Yes. You can try to mitigate the risk a bit, but it's still there. Just think of all the other firmware you can update, not just UEFI.
Also, for large contracts, rented servers are usually bought new and not re-used for multiple reasons (this is why providers have used server markets!).
Virtualization overhead is so low nowadays that it's not worth the risk. Provisioning is much harder, too.
The entire Finance industry would like to disagree with you.
On the other hand, you don't run latency sensitive code in the cloud, you run on your own hardware. Since GCE is a cloud, no reason not to have the virtualization.
The Google announcement didn't move their stock much at all.
From the post:
> Google Cloud will offer AMD FirePro S9300 x2 that supports powerful, GPU-based remote workstations. We'll also offer NVIDIA® Tesla® P100 and K80 GPUs for deep learning, AI and HPC applications that require powerful computation and analysis. GPUs are offered in passthrough mode to provide bare metal performance. Up to 8 GPU dies can be attached per VM instance including custom machine types.
Disclosure: I work on Google Cloud and pitched in on this.
What kind of asshole move is this? Why not just say "here, you can use it now, good luck".
> Tell us about your GPU computing requirements and sign up to be notified about GPU-related announcements using this survey. Additional information is available on our webpage.
Survey: https://goo.gl/mgEI9X Landing Page: https://cloud.google.com/gpu/ (which also links to the survey and indicates as you rightly presumed, that if you fill it in, we'll add you to our waitlist)
Disclosure: I work on Google Cloud.