GPUs Are Now Available for Google Compute Engine and Cloud Machine Learning
cloudplatform.googleblog.com
cloudplatform.googleblog.com
As I said back in November [1] for our initial announcement, the most exciting thing (for me) is that we let you mix and match cores and GPUs. You can see that spelled out in a nice table in the docs [2].
[1] https://cloudplatform.googleblog.com/2016/11/announcing-GPUs...
There are also perpetual free tiers for several GCP services - AppEngine, BigQuery, Firebase.
No comments on your specific feedback, but please do join GCP NEXT March 8-10 for some exciting updates [0].
(Work on Google Cloud, NOT in marketing :))
(I work on it.)
Since there's a dollar cap as well, I don't see the problem with extending the trial out to a long calendar time-period.
I'm designing the distributed backend for my Deep Learning startup now (https://signalbox.ai), this is going to be great. Awesome work!
Cpu costs may be something 5 to 10 times more expensive in the cloud, but network are close to 100 times. Any hosting provider will offer you a 250 mbps unlimited network access for your macine, whereas consumming that much bandwidth in google cloud for a month will cost you more than 1000$.
How come is the difference that big ?
> Cpu costs may be something 5 to 10 times more expensive in the cloud
So are CPUs more or less expensive in the cloud?
Also I assume by "cloud offerings" you mean AWS, Azure, Google Cloud and the like, and by "regular hosting providers" you mean GoDaddy, BlueHost and the like? Or perhaps you mean Paas vs Saas offerings?
If you are interested in what 1/100th of AWS price looks, have a look at Online's server order page for example [1].
Most of the providers like Hetzner/OVH provide former, while GCE provides the latter. I'm not saying it's bad, in fact for most of the people 1Gbps would be more than enough. But it's not something that is fair to omit.
Disclaimer: I don't work for Google, just from my experience.
I am sure Hetzner would be able to offer the same as an add-on if you contacted them directly.
That being said, you would continue to pay orders of magnitude less for your bandwidth then you would at Google or AWS. There's a bubble in cloud bandwidth pricing, and I don't think it's value-related.
What i don't understand is the network cost factor of x100 to x1000
To a first approximation, CPU and RAM are required to get your site up and running at all. Bandwidth is less so. Bandwidth scales up as your growth scales up.
So it makes sense for cloud providers to make CPU and RAM relatively cheap, and charge unreasonable prices for bandwidth. If you're growing, you're more inclined to pay, since you're seeing success. Plus you're already locked in at that point.
The bigger challenge for large ML models is the memory you'd need to back it. But with GCE you can happily do a custom machine with up to 6.5 GB per vCPU, just so you can fit the output ;).
I see one of our projects has gotten a quota of 16 GPUs in asia-east1, us-central1, us-east1. However we seem to have been allocated nothing in europe-west1. Is this an error, or simply something I have to manually ask for somewhere?
"Quota 'NVIDIA_K80_GPUS' exceeded. Limit: 0.0" :(
I'm so psyched to finally be able to use GPUs on GCE instead of AWS, so any help would be appreciated.
We do not currently apply sustained-use discounts to GPUs, but we may do so in the future (we need to gather data first on usage, to understand if people will be running 24x7) [1]:
> You cannot attach GPUs to preemptible instances. GPUs do not receive sustained use discounts.
1. Are GPUs covered by the free trial (ie, can that $300 be spent towards GPU instances)?
2. How is support for GCP?
I've been curious about trying GCP, but held off over GPU support (since AWS was covering my needs and I do ML stuff mostly) and general support (since Google doesn't have a great reputation for supporting products).
Also, perhaps an affiliated person can chime in with something about the roadmap to stable GPU support. (It currently says there may be breaking changes.)
2. Unlike consumer-facing products, GCP is focused on business. We offer paid support plans [1] with high (measured) customer satisfaction. I know several of the people in the support teams (and you see them here on HN as well), and we're really trying to defeat the meme of "Google doesn't do support".
As far as "stable GPU support", this is just confusing language surrounding our usual "Beta" terms [2]. Once it becomes Generally Available, no changes would be made. But moreover (for GCE anyway), we don't make API breaking changes from Beta to GA (Beta to GA is "just" about stability in production).
What do I have to upgrade to? Isn't GCE all per-minute? Do I just have to pay $0.05 for a standard instance and then use the $300 from my trial?
Also, is coin mining against any TOS? I'm not planning on doing it, I'm just curious.
Coin mining is not against the TOS. However, because it's usually economically irrational, it's usually abuse. If you don't pay for your GPUs (fake credit card) it's awfully economically rational though ;).
http://web.archive.org/web/20160704054747/http://www.complet...
So, around 4 cents an hour?
"The board is designed for a maximum input power consumption of 300 W"
So one hour at home:
Power used: 300 Wh
Price: $0.20/kWh
Total cost: $0.06/hour
Google cost: $0.70
You could say it's 10X more expensive, but you need to include all additional costs, power is just a fraction of it. The GPU itself is $5k, which over - let's say - 3 years would cost $5/day, or $0.20/hour, and if you're using it a third of the time, just that amounts to $0.60/hour. Add everything else (you might use less than 1/3 of the time) and it would be far more expensive than using it in the cloud. As it is expected to be.
But the double precision perf of M60 is very low and less than the K80 due to limitations of the Maxwell microarchitecture. So we could say K80 is the last decent dual-GPU Nvidia card...
Is there a GCP Image with CUDA drivers pre-installed? Or is that not possible with the hardware architecture?
But IANAL.
Edit: still requires nvidia-docker, or a hand crafted docker command that replicates nvidia-docker.
[1] https://docs.blender.org/manual/en/dev/render/cycles/gpu_ren...
You might find good fun with our Preemptible VMs, which I think will land about ~$7/month: https://cloud.google.com/compute/docs/instances/preemptible
If you manage to build you app/infrastructure in away that can survive nodes shutting down at random times (or if you don't care about restarts) then you can reduce your infra costs for 80%.
We have a staging cluster running on preemptive instances and as soon as one instance goes away we get a different one. Everything gets deployed automatically. Regular internal users checking out various webpages don't even notice.
We're looking into changing our 24/7 infra (which needs to be 24/7) to something that can be run mostly on preemptive (with a couple of normal instances for services that can't be randomly killed).
Super happy about our move to GCP and our K8S experience.
Otherwise, if it'd load an AMI - ideal!
Does anyone have, specifically for AI/ML, a list of "If you want to do X, use Y" for Google Cloud offerings? The official list (https://cloud.google.com/products/machine-learning/) doesn't help much. Would appreciate if it's explained by layer (higher such as ready-to-use Speech Recognition and lower where you possibly need to setup some infra stuff).
EDIT: I'm looking something like this (explains AWS offering by layer) but for Google -> https://aws.amazon.com/blogs/ai/welcome-to-the-new-aws-ai-bl...
https://hackernoon.com/what-are-the-google-cloud-platform-gc...
Or for the really lazy:
https://twitter.com/gregsramblings/status/832223967096090624...
http://austriantribune.com/informationen/165669-googles-new-...
https://medium.com/@thetinot/google-clouds-spot-instances-wi...
> You cannot attach GPUs to preemptible instances. GPUs do not receive sustained use discounts.
So sorry, no.
I think on P2, Amazon's cards are remote, and there is pretty significant latency when using them for time-sensitive computing.
What kind of performance can we expect from GCP compared to AWS P2?
I'd love to see your findings, I'm curious as well!
Separately, can tensorflow models trained on Cloud ML be downloaded yet?
[1] https://www.tensorflow.org/versions/r0.11/how_tos/meta_graph...
The copy on https://cloud.google.com/ml/ under Portable Models says "In future phases, models trained using Cloud Machine Learning can be downloaded for local execution" so I hadn't looked into it further.
[Edit: Okay, we've decided that the exported model it produces is what you'd expect. We're going to update the landing page, once we can agree on what it should say.]
from cloud.google.com/gpu (our landing page). That's currently limited on availability of hardware, testing, etc. but it really should be "soon". Note though that P100s are massive and expensive, so we don't intend to get rid of K80s or anything once we have P100s.
Always interested in what kinds of things one can do when new offerings like this are made.
GPU is powerful because a GPU usually has 100+ cores. Each core is weak and inefficient, but power adds up when you have 100+ cores available.
https://en.wikipedia.org/wiki/General-purpose_computing_on_g...
Having said that, on-premise or bare-metal is about 25% faster than AWS P2 for large datasets (over 2TB). For smaller datasets that may fit in-memory, they function about the same.
and you have to use a remote windowing dealie
(I'm joking, but there has to be a day when it's possible, right?)
https://lg.io/2015/07/05/revised-and-much-faster-run-your-ow...
(We have support, more than 600 security engineers alone, we just offer commercial support for a commercial product.)
Is where you can talk to a human every time you file a ticket related to Google Cloud Platform.
(disclosure: a human that answers those tickets)
While a number of people push that false meme repeatedly, everything I've heard from people who've used it (or worked on it, but the latter comes with obvious bias) is that the paid human support on GCP is good.
Heck, on the consumer side, I've gotten good (quality and speed) human support on Google Express, too.
Honestly though I can't say I've ever really needed support on GCP, most of the times I've raised tickets its been due to funny behaviour (slow spinning disks in us-central1 was the last one).
Source: Pay for many Google services, received support when needed.