New P2 Instance Type for Amazon EC2 – Up to 16 GPUs
aws.amazon.com
aws.amazon.com
Just as an example of the change this entails for deep learning, the recent "Exploring the Limits of Language Modeling"[1] paper from Google used 32 K40 GPUs. While the K40 / K80 are not the most recent generation of GPU, they're still powerful beasts, and finding a large number of them set up well is a challenge for most.
In only 2 hours, their model beat previous state of the art results. Their new best result was achieved after three weeks of compute.
With two assumptions, that a K80 is approximately 2 x K40 and that you could run the model with similar efficiency, that means you can beat previous state of the art for ~$28.8 and could replicate that paper's state of the art for ~$7257.6 - all using a single P2 instance.
While the latter number is extreme, the former isn't. It's expensive but still opens up so many opportunities. Everything from hobbyists competing on Kaggle competitions to that tiny division inside a big company that would never be able to provision GPU access otherwise - and of course the startup inbetween.
* I'm not even going to try to compare the old Amazon GPU instances to the new one as they're not even in the same ballpark. They have far less memory and don't support many of the features required for efficient use of modern deep learning frameworks.
And at the end of the month you still have a modern GPU to play video games on etc.
Of course if you have money to burn, these are great.
I would assume they have (or have access to) a computer fitted out in a research lab or something which would be seperate from their "gaming rig".
If you are into the topic and want to work on things outside of assigned projects or course work, upgrading a personal machine doesn't seem far-fetched if you can make room in the budget for it. $700 isn't cheap, but doable for many as a one-off purchase if they want it.
Those who are predisposed to buying & installing high-end graphics cards are likely to be either gamers, interested in machine learning or both
While I'm sure your statement is true for persons, I doubt its true for the computers themselves.
Most people I know who got into the field via gaming were either completely incompetent and have since left, or do not label themselves as 'gamers' anymore.
Who are the majority of machine learning researchers? PhD students in CS (aka majority male) between 22 and, what, 26 years?
What percentage of male computer science students between 22 and 26 years DOESN'T play games?
I did it a few months ago with a TensorFlow model I was training. Configured it for free then all the training time was fully used after configuration.
I assume it'll be the same with these new instances. I have a backlog of stuff I wanted to train but had no GPU to train it on (I travel a lot so couldn't buy a desktop, and was almost biting the bullet to buy one of those monster laptops with dual 1080's, although really I was after at least 12GB of VRAM so this is fantastic news for me)
Also, consider the efficiency of owning such a 16 GPU beast: if you don't have tasks to run 100% of the time, then you're losing a lot by just keeping it idle. Cloud providers can reuse the same GPU for other clients in the meantime.
Surely there are some cases when you are better off computing on your own GPU, but I think in all of them you assign pretty low value to the time spent.
[1] https://buildazure.com/2016/08/09/azure-n-series-vms-and-nvi...
Now I'm using a P2 instance to see how much faster it runs the same neural net. I set the max price to $0.50 in a spot price bid and got it immediately. I'm not sure how to find out exactly how much it is being charged however.
The 1.1km forecast runs on 144 GPUs, the 2.2km probabilistic ensemble forecast is computed on 168 GPUs (8 GPUs or 1/2 node per ensemble member). The 7km EU forecast is run on 8 GPUs as well.
Here's the scientific description of the model: http://www2.cosmo-model.org/content/model/documentation/core...
Google Cloud Platform just released Cloud ML beta with different pricing model, see https://cloud.google.com/ml/
Cloud ML costing $0.49/hour to $36.75/hour, compared to AWS $0.900/hour to $14.400/hour
The huge different of $36.75/hour (Google) compared to $14.400/hour (AWS) make me wonder what Cloud ML are using, they mentioned GPU (TPU?) but not exact GPU model.
I would say that while I still have to evaluate Google Cloud ML, AWS offers a far far better experience and cost effectiveness.
I couldn't believe how much better AWS is for running TensorFlow models than Google Cloud.
(work at GCP)
Thanks for replying.
All of the instances are powered by an AWS-Specific version of Intel’s Broadwell processor, running at 2.7 GHz.
Does anyone have any more information about this? Are the chips fabricated separately or is it microcode differences?I think this is the right talk by him, but there's one where he implies some differences. For instance, that because they can promise exact temperature ranges they can clock them differently. My guess is same fab, just binned differently, or maybe different packaging.
Not likely.
The AWS CG1 instance type was introduced in 2010 (NVIDIA Tesla C2050 GPU). The G2 instance type was introduced in 2013 (NVIDIA GRID K520 GPU). Now the P2 instance type has been introduced in 2016 (NVIDIA Tesla K80 GPU). At this rate, we probably won't see another GPU fleet in AWS until 2019.
No. But in my view had this been written by an independent journalist, they would describe it as "TensorFlow by Google". The project is virtually completely driven by Google, so "By Google" is a more distinguishing classification than "Open Source". We can speculate but strong impression is that Amazon doesn't want to give kudus to Google.
"We hope to build an active open source community that drives the future of this library, both by providing feedback and by actively contributing to the source code."
(This might be an issue of my account though - having had only small bills so far)
They're likely doing it as a way to control demand spikes because they have such a smaller pool of GPU instances to go around. Also, it prevents Joe Schmo from wasting one to run a web server because he doesn't understand performance and "just picked the best one".
My second thought: "I wonder if Amazon use their 'idle' capacity to mine cryptocurrency?"
With respect to my second thought, at their scale, and at the cost they'd be paying for electricity, it could quite possibly be a good hedge.
But most of these markets are very small compared to Bitcoin.
Mine other cryptocurrencies that are profitable, or may become profitable. It is VERY, as in EXTREMELY, profitable to mine large stakes in cryptocurrencies that the market hasn't discovered yet.
Especially ones that aren't even listed on an exchange yet and mining is the only way to earn it.
The Darkcoin (now renamed to the politically non-threatening name DASH) instamine was a notoriously profitable use of GPU clusters on AWS.
Also, if you are paying to mine currencies that don't have an exchange value yet, you are just speculating. The mining isn't profitable at that point because the currency has no value. It's just as likely that you chose to mine a dud. It's exactly the same thing as opening a copper mine at a loss and hoarding the concentrate hoping that it will gain value.
you say that like its a bad thing
it is a very profitable endeavor
But on the other hand, these machines aren't just sitting there powered off when they aren't provisioned. They have to keep them warm so they provision ultra quick.
So maybe there is some level of mining they could accomplish with their "idle" cycles, given the efficiency of the datacenter, the power of their GPUs, and the price of Bitcoin, which would be profitable. I don't know.
But they don't even have to be making a raw profit vs the electricity they are consuming. If, as you say, the spare capacity must be kept warm anyway, then they only need to be profitable vs the delta of electricity consumption between idle and fans blowing 100%.
So whatever capacity is left beyond the fixed and spot market _could_ still be earning them something.
[1] http://cryptomining-blog.com/6090-crypto-coins-to-check-out-...
Edit: fixed spelling.
But at the price they're charging, if demand exists, they just call nVidia and order another batch of cards
After a few months they can estimate the demand and increase supply accordingly (iff profitable)
Also, they'll get a lot of instances up and running before publishing the post. At the moment, p2 spot prices seem very low (~9c for p2.xlarge).
Especially in the GPU space, a lot of calculations will likely not be very time-sensitive. The spot market is optimal to control demand and ensure high utilization.
Launch Failed We currently do not have sufficient p2.16xlarge capacity in zones with support for 'gp2' volumes. Our system will be working on provisioning additional capacity.
Chips: 2× GK210
Thread processors: 4992 (total)
Base clock: 560 MHz
Max Boost: 875 MHz
Memory Size: 2× 12288
Clock: 5000
Bus type: GDDR5
Bus width: 2× 384
Bandwidth: 2× 240 GB/s
Single precision: 5591–8736 GFLOPS (MAD or FMA)
Double precision: 1864–2912 GFLOPS (FMA)
CUDA compute ability: 3.7
Is that a good deal for $1/hour? (I'm not sure if a p2.large instance corresponds to use of one K80 or half of it)How much would it cost to "train" ImageNet using such instances? Or perhaps another standard DDN task for which the data is openly available?
______
Comes to just under $50,000 for the server or roughly 4.5 months @ $14.40
If you do absolutely care about not being multitennanted, you can stipulate exclusive hw allocation when requesting the VM.
Well, an attempt is made. As we can see in the case of WebGL, there are a multitude of interesting effects.
Disclosure: I work on Google Cloud.
From 0.90-14.40 USD/hr.
I work with live transcode, and it can be beneficial to run on g2 instances. Running c4.4xlarge I can transcode a good number of 1080@30 in with 1080/720/480/360/240@30 out. With a proper g2 instance I can transcode more simultaneously.
So cost efficiency really depends on sustained traffic levels. I scale out currently using haproxy and custom code that monitors my pool and scales appropriately. But I monitor sustained traffic levels to know when it makes financial sense to scale up.
If your main concern is transcode speed, CPU is likely sufficient -- I am unable to transcode faster than real time with live transcode.
Well, yeah. Maybe quantum computing will change that one day!
What is it you're working on? Colour me intrigued.
All in all, I look for somewhere around 0.25s to complete the rest of the work and deliver the segment to CDN.