Which GPU(s) to Get for Deep Learning
timdettmers.com
timdettmers.com
I saw a lot of "should we use Cloud, no its crazy a GPU only costs $X". The key is that if you believe GPUs are going to get updated every year, and/or the best thing for ML may change (see TPU and plenty of startups with custom hardware) then suddenly buying hardware for 24? 36? months isn't as obvious.
We (and AWS and Microsoft) have K80s because Maxwell wasn't a sufficiently friendly all around part. We're all going to offer Pascal P100s and in the future V100s. The challenge for buying your own is that P100s are available now-ish and V100s may be available in less than 12 months.
Buying a P100 or similar part today doesn't mean it won't still be working in a year, but it will suddenly mean you've bought a part that now has much worse !/$ in just N months. If you have an accounting team that is spreading your $Xk over those 36 months, the reality is that you have two options: tell everyone they have to make use the old parts ("we're not buying new GPUs until it's been 36 months!") or realize you're going to get a lot less use out of them.
To be clear, the progress in this space is really impressive. And the same problem above applies to us (the cloud providers). Despite my obvious bias, if I were fired today, I'd be renting to do deep learning just based on the roadmap alone (not to mention the ability to suddenly spin up and down).
Again, Disclosure: I work on Google Cloud and want to sell you things that train ML models :).
I don't disagree that a consumer board is about that price, but they're not apples to apples. (Either in memory size, reliability or both). I'm fine with that being the real complaint: (major) cloud providers only sell the Tesla class boards, and they're really expensive ;).
The real attractiveness for cloud at this point is if you are going to train your model with 8-GPU or more, that is likely not feasible for individual enthusiasts, but demand for such machine is rare for hobbyists anyway.
There are also other hidden costs that aren't factored into that number. I'm having a hard time getting anything reasonable for under $750/month for 24/7 usage: https://cloud.google.com/products/calculator
https://arxiv.org/pdf/1704.01444.pdf
They say they used four Titan X cards for a month. I was actually looking at google cloud machine learning since I thought being targeted to tensorflow would be the most cost effective. But it's 0.49/hour per ML training unit, or 1.47/hour per GPU. A basic gpu gives 3 training units, so I thought that something approximately similar would require 12 training units, which comes to something like $4000+ a month. Maybe I completely misunderstood the resources being offered though, because you are right that cloud engine costs seem much lower.
I'll have to revisit the math here, though it worries me that it's not at all clear that a K80 will be much faster than a Titan X for a given problem. E.g. https://www.amax.com/blog/?p=907. It would also be really nice to get some pricing and benchmarks for the new TPUs, assuming they are priced better.
Maybe part of the problem is that vendors are not making it remotely easy to even understand what performance you'll actually get for a given price.
Payback time according to my own calculations depending on GPU model and machine used compared to the various cloud offerings is between 4 and 6 months when used continuously.
It would be nice to see Nvidia or someone expand on this, so that users who have to make this choice can do so without guessing. If Google or AWS or M$ could publish reliability information, that'd be cool too.
Illustrative case: I run Monte Carlo work on GPUs and administer a local compute cluster. I tested a workload on a 16 GB P100 and a GTX 1080. A 12 GB P100 costs (academic, EU) 5000 euros while the latter costs 700 euros, but the performance difference is about 2x. Still, when we ask Nvidia reps, they say not to bother installing GTX cards in our cluster, because they aren't designed for 24/7 work, not commercializable etc. Even so, the GTX would have to burn out 3 times before the choice of P100 breaks even. Burning out three times means GTX 1080, then 1180, 1280 etc.
It's cheaper, allows gradual adoption of deep learning and is a familiar toolset for folks already doing some sort of machine learning.
I won't comment on research (not our domain). Depreciation on hardware is pretty standard - that being said it also comes with established SKUs from dell,hp,cisco,.. with proper support.
Analytics clusters (while hard to manage) are fairly robust already to job failures. The cost savings just makes a ton more sense when you are doing continuous workloads for different use cases.
If so, that's fine-ish, but then people have to wait (you're either full and people are waiting or you're at less than 100%). My preference is to run for XX minutes per job on-demand (per person). If you have tons of non-overlapping users, you can absolutely aggregate them. But how many do you buy and how quickly do you upgrade to newer hardware?
[1] https://cloud.google.com/dataproc/docs/concepts/gpu-clusters
That being said - it can usually be every 6 months with a renewal of every 2 years or so. Your hardware timelines aren't far off. That being said - I was getting that upgrade cycles don't matter as much.
GPU clusters are enough of a new thing for enterprise yet that the jobs being run aren't even that high spec yet.
My opinion nothing more not claiming this is fact: Google has marched so far ahead of the rest of the world they aren't really paying attention to where current clusters and usage are. It seems a lot of the cloud usage is oriented towards startups and researchers (which isn't a bad thing, most folks in DL are researchers).
For enterprise, they might offload some workloads. There are definitely some workloads where cloud resources (spin up and shut down) make a ton of sense. Cloud servers are overly expensive otherwise.
By the way, I appreciate the useful comments you've left in the past on your experience for gce.
(I have no association with Floyd)
I have a prepared draft for a blog post exactly on the topic of cloud computing and deep learning. I did not finish it as I thought that there would not be much interest in the overall question since most people will just buy GTX cards. However, it seems that there is quite a confusion going on what makes sense and what does given certain circumstances. I think I will finish that blog post now and post it in the next days.
If you want me to discuss certain questions regarding deep learning hardware and cloud computing let me know here.
They provide a dedicated server with a GTX 1080 for ~99e/month (111$/month) with adequate CPU (i7-6700), 64G Memory, 500G disk space and 50TB Traffic - there are also on-demand offerings by GCP and AWS, but I do not think they can match the offer by Hetzner: https://www.hetzner.de/us/hosting/produkte_rootserver/ex51ss.... Keep in mind that I am talking about R&D, in particular training of networks - which has lower expectations on availability than e.g. later inference.
Disclosure: I am not affiliated with Hetzner and did not test it yet, but I have a ~40e dedicated machine there for occasional number crunching and everything worked so far (no availability or hardware issues).
Furthermore I am not sure about the exact specifications (Memory-Size of GPU) or if there are different types of 1080s that differ significantly in Deep Learning Performance (see also: https://github.com/jcjohnson/cnn-benchmarks).
0.5 * 24 * 30 * (30 c/kWH) = $108
However, it'd be a third of that in the states (may be cheaper in WA).
http://hexus.net/media/uploaded/2017/5/30f5633b-1bbf-49b7-9f...
If AMD provided good implementations of similar libraries with OpenCL and shipped them with their drivers (similarily to Nvidia's CUDA toolkit) that would be much better than supporting CUDA, but leaving the ecosystem bare bones again...
What's more interesting is the disruptive change to the underlying programming model for Volta (thread-within-thread) that will break a lot of existing CUDA code without refactoring. So if there's no GeForce Volta option, I suspect a lot of that code won't get refactored. So Vega just needs to not suck and to deliver decent TensorFlow/Torch/Caffe performance to succeed as an alternative to NVIDIAopoly.
Tesla P100 was the first real HW-level divergence with its 2x FP16 support. But because we still can't have nice things, GTX 1080 was the first GPU with fast INT8/INT16 instructions, followed by the mostly identical except much more expensive Tesla P40. So we ended up with the marchitecture nonsense that P100 was for training as P40 is for inference despite being mostly identical except as noted above.
I'll assume Volta unifies INT8/INT16/FP16? And I think it's OK if the Tesla card has higher tensor core performance, but if the tensor core on GeForce is slower than its native FP16 support, I can only conclude NVIDIA now hates its own developers and has decided to sniff its own exhaust pipe. Isn't having to refactor all existing warp-level code for thread-within-thread enough complication for one GPU generation?
Also, if consumer Volta ends up with craptastic FP16 support (ala 1/64 perf in GP102 vs GP100, slower than emulating it with FP16 loads and FP32 math), NVIDIA will create a genuine opening for AMD to be the other GPU provider in deep learning.
If you can't do what Blender did and write everything from scratch including your primatives you'll be much slower than CUDA.
But there is also a cost to it the Blender Cycles OpenCL code is nearly 5 times as big as their CUDA code and Cycles on OpenCL is still not at a feature parity with CUDA.
The change in branding would likely not yield much for it.
On the other hand, I think you're underestimating the fact that OpenCL is an open spec and therefore has support from the FOSS world. CUDA has always been criticized for being closed source.
Even though OpenCL doesn't get much praise, and actually get criticized a bit (I'm personally not a big fan because writing OpenCL code is like bending over backwards). It has a lot of potential with new libraries coming up[1] which can possibly make it atleast on par with CUDA. Once that happens, AMD's cost effective cards will put it in the race.
It still relies on many pretty old libraries with questionable performance at best.
The newer stuff is good but it's tailored to AMD GPUs, don't expect interoperability when it comes to any reasonable performance if you want to run ROCm OpenCL on NVIDIA or Intel hardware.
Many of the core libraries are not exactly open but these are primitives.
If say cuDNN would become open source today there won't be any benefit from it, you won't be able to make it run faster than what NVIDIA has already achieved.
As for the vendor lock well this is tricky, AMD also tried to vendor lock their initial compute APIs they've fallen back on OpenCL because it wasn't working.
ROCm isn't exactly vendor neutral while it can run on nearly any OpenCL compatible hardware it's tailored and optimized for AMD GPUs which means it would run like utter garbage on anything else.
It also supports interoperability with CUDA via the HIP compiler.
And in all honest at least for scientific computing OpenACC (https://www.openacc.org/) might make this argument irrelevant at least as far as general use code protability goes.
If you have any question, please ask!
While AMD doesn't create something like cuDNN to go along with their cards, no real work will start being done in porting most important DL libraries to AMD cards. And even in that case, it will be a lost generation for AMD, only in the 2nd generation where AMD actually offers a real alternative in DL to NVIDIA, you will start to see some real progress being done with it.
They also stated this last year:
> On top of ROCm, deep-learning developers will soon have the opportunity to use a new open-source library of deep learning functions called MIOpen that AMD intends to release in the first quarter of next year. This library offers a range of functions pre-optimized for execution on Radeon Instinct cards, like convolution, pooling, activation, normalization, and tensor operations
http://techreport.com/review/31093/amd-opens-up-machine-lear...
Their SpecPerf View benchmarks were a total FUD they compared NVIDIAs consumer (GeForce) drivers against Radeon Pro drivers.
The SPV benchmarks look impressive until you realize they are lower than a Quadro M5000 which is based on the same Maxwell chip that drives the 980ti.
NVIDIA's consumer drivers (and the non Pro AMD ones) are simply horrible for CAD and other professional workloads.
AMD has been touting better performance than NVIDIA GPUs in compute since Fiji and they do not deliver.
It's also important to note that VEGA will have FP64 performance at 1/32 while the P100 is at 1/2.
I'm sure you can find an OpenCL workload that NVIDIA would be utterly trash in because the code is not optimized heck working with 256 batches as recommended for AMD rather than 1024 as recommended for NVIDIA would achive that alone.
If that GPU is a real bottleneck for you, then you're much better off spending money on GCP/AWS's GPU offerings. That's because consumer GPUs get superseeded every year and online offering's price will only go down.
So you can spend 10% of $1k every year and keep getting better return on compute / dollar every year.
[1]:https://www.newegg.com/Product/Product.aspx?Item=N82E1681448...
GCP and AWS have old GPUs and they are really really expensive. If you expect to run workloads for a long time, it would be more cost efficient to buy your own hardware.
The 1050 is a beginner card is it is perfectly fine to learn and run small nets. More importantly, you can decide if machine learning is for you. Then, comes to second investment, which is actually running real-world models.
Although the online GPU offering is expensive (you can also look around for cloud GPUs with lower SLA requirements for a lot cheaper) you'd be using them for a lot lesser time.
Even if you go with the big three, you can get good pricing if you look around. AWS has a single GPU with 1,536 CUDA cores and 4GB RAM. Looking at the Spot Pricing[1] which is $0.25 an hour, you can get 3000 hours of compute for $750, which is the price of 1080Ti (which is the most cost effective card on the market today).
Now, I would say, 3000 hours is more than enough time one needs to run whatever they want to. You can get a lot more if you go with some other service other than the big three.
If your usage exceeds that, then you are probably using it for commercial purposes, in which case you need to account for a lot more variables (downtime, maintenance, etc). Then you should also consider the deprecating perf/$ your GPU gives you every year over the new ones in the market - especially since the perf jump is a lot more than what we're seeing with CPUs.
If you are doing it for learning, I'd say you're doing something wrong, because you shouldn't need that much power. Also, consider the electricity costs and possible over usage of your computer.
[1]:https://aws.amazon.com/ec2/spot/pricing/ (check US East N. Virginia)
* Setting up machine on AWS is more complicated than locally, and requires some admin skills.
* If you use spot instances, you need to handle checkpointing, which requires persistent storage, and all of this stuff requires even more admin skills
The goal of a person who starts working with deep learning is to learn deep learning, not how to setup machines, manage them, work with checkpoints, etc.
Also, don't forget that there's a large market for used GPUs and you can get real bargains.
There are tons of guides online where you can learn how to do so in <5 mins. Setting up the computer by yourself is more complicated I would say, and would take a first timer days rather than hours.
>If you use spot instances, you need to handle checkpointing
Again, not a big deal to learn.
>The goal of a person who starts working with deep learning is to learn deep learning, not how to setup machines, manage them, work with checkpoints, etc.
I mean, if you're buying a GPU and setting it up, you're more than likely assembling the computer by yourself. You'll also have to maintain it properly. Then you'll have to look for correct drivers and other software which can get frustrating (it did for me).
On the other hand, I just could use a step by step process for the AWS instances since they had a few specific types of GPUs and I didn't have to even think, just copy/paste the commands from the webpage to the terminal. There are even AMIs which setup everything for you, which would require even less effort, but I don't trust them so I go with a clean disk.
Moreover, learning how to use AWS is a much more valuable skill than putting together computers so time well invested I would say.
http://webcache.googleusercontent.com/search?q=cache:04az_uB...
There is no reason, not to get a Pascal GPU at this point, the performance is simply superior. For a lot of models, we are talking about days of training time, so 20%-30% time saving is significant, and we are not even touching the part that large VRAM enables bigger batch size which you won't get in middle/low end GPUs.
They're fairly easy to get now too, in the beginning it was rather hard to get them.
http://vfio.blogspot.com.au/2014/08/vfiovga-faq.html https://www.reddit.com/r/linux/comments/2twq7q/nvidia_appare... https://www.redhat.com/archives/libvirt-users/2014-October/m...
I just quickly googled this so there are probably better sources. Some of these are old but I can tell you firsthand that this is still the case.
There are workarounds for certain hypervisors (KVM mainly) but it's very unlikely that this would be deployed in a production environment.
Unless you have an unlimited supply of free electricity and don't care about the increased hardware management overhead, it's a waste of money to buy Pascal GPUs for large-scale deep learning.
The following cards have much more optimized deep learning silicon and are publicly available /right now/:
- Nvidia Tesla V100 (Tensor cores only: 120 TFLOPS FP16)
- Google TPU2 (180 TFLOPS FP16)
Additionally, Intel Xeon chips with the Nervana deep learning accelerator built-in will probably be available early next year.
If you must control the physical hardware yourself and can't use cloud services, go buy Tesla V100s or wait for the Nervana Xeons.
https://www.nvidia.com/en-us/data-center/volta-gpu-architect...
'Tensor Processing Unit' has become some what a phrase used to confuse (trick?) people into thinking it's some new type of processing, but its not. If the GPU says it has a Tensor Processing Unit, it just means it operates with a lower precision, but you have to remember that the overall power consumption of the GPU doesn't really change at full use compared to using higher precision chips. So you're actually missing out on the cost effectiveness of using lower precision.
If anything the 'Tensor Processing Units' just take up unnecessary die space when included high higher precision units because they're bad for training compared to higher precision compute units.
I realize they are physical parts of the die, which is why I said in my last sentence about how they take up unnecessary die space. Why? Because that die space will be better used for higher precision FP units which will be useful in training, which is more important than inference for most of the people in this thread.
If you are asking about using the iGPU for a bit more power, most frameworks don't support them. The only frameworks with which you can use with iGPUs are which have OpenCL support, like caffe[1] (even for those you need to have a recent enough CPU)
Tensorflow either uses CPU or CUDA (Not even AMD, but the support is coming for openCL)
With less than 1% error. Here's a link to the paper: http://yann.lecun.com/exdb/publis/pdf/lecun-01a.pdf
1060, 1070, or 1080Ti are all good choices, depending on your budget and ambitions.
zeros: each element of this array is a picture of a "0" ones: each element is a picture of a "1" ... nines: each element is a "9"
Compute the average of each array: the "average" 0, the average 1, and so on. Then classify new digits based on which average element they are closest to, using Euclidean distance. This sounds way too dumb to do any good, but the MNIST digits are normalized so well that this actually does something.
Even better: you can have a neural "network" that has zero hidden layers. This actually achieves almost respectable performance, believe it or not.
On MNIST I think you get something like 60% accuracy.
https://lambdal.com/deep-learning-devbox
On average it takes a SWE or DL engineer a few days to set up a unit from scratch. Your company probably burns over $2,000/day so every day your DL engineer or SWE isn't up and running costs you money.
The anti-competitiveness is getting worse too. Charter just took over TWC in my area.
Assuming 2.5MB/s Upload capacity with local data, you've uploaded it to your deep learning machine in half a day (100000 / 2.5 / 3600 ~ 11 hours) - which is not that much, as most of your time will be used for development and fine-tuning of your deep learning tool chain anyway. In most cases the data is accessible via a public-facing service, and assuming 1GBit/s bandwidth you've downloaded 100G in 13 minutes (100000 / 125 / 60 ~ 13 minutes).
I know egress is one of the more expensive cloud services (e.g. compared to compute and storgage) at AWS, GCP, etc., but if I upload data to my learning system that's ingress AFAIK which is mostly free or less expensive. Btw. current Egress is like 0.1$/GB, so 100G ~ 10$.
Don't get me wrong, I am not saying you should always train in the cloud, but I do not think slower Upload or Ingress are the limiting factor.
1. Can laptops with, say NVidia 1070 or 1080 GPUs keep them cool at their stock frequencies for a few hours? My work laptop with non-U i7 (Thinkpad T440p) starts thermal throttling in just a few minutes when I compile something large.
2. Wouldn't two NVidia 1060 or 1070 outperform a single 1080 for training, assuming batch size is kept low enough so that each batch fits in a single card's memory?
Is there anything better now than the GTX 1080 Ti?
At the NVIDIA conference all the second tier hosting companies were promoting the P100 (at NVIDIAs insistence) but when pressured admitted that their big customers now deploy 1080TIs. Paying for P100s is sort of a clown move even if you're spending someone else's money. The P100 starts at like 10x the price of the 1080TI and isn't much faster and again has just 12GB of RAM (16GB if you pay $4000 more, ludicrous)
Also if there do build with multiple gpus, watch out that you can give both cards a full 16 pci lanes, not a given in a lot of motherboards.
nvidia-smi
Tue May 23 01:41:58 2017
+-----------------------------------------------------------------------------+
| NVIDIA-SMI 375.39 Driver Version: 375.39 |
|-------------------------------+----------------------+----------------------+
| GPU Name Persistence-M| Bus-Id Disp.A | Volatile Uncorr. ECC |
| Fan Temp Perf Pwr:Usage/Cap| Memory-Usage | GPU-Util Compute M. |
|===============================+======================+======================|
| 0 Graphics Device Off | 0000:01:00.0 On | N/A |
| 57% 84C P2 257W / 250W | 11010MiB / 11171MiB | 95% Default |
+-------------------------------+----------------------+----------------------+
That's during training. I'm running a minimal desktop to have maximum ram for the minibatch (and it still isn't enough but it will have to do).We found that the GeForces tend to burn out when under heavy load, whereas we've not had a single Tesla series ever burn out.
We've been using Titan-X and 1080 Maxwells in some Broadberry 4u chassis for the last year/18mths and we've had no burnouts so far.
I'm buying replacement pascals, and I can't justify teslas when I can get 8 geforces for the price of 1...
On some servers we introduced the GTX 1080, either along-side a K40/K80 or two per chassis (see http://arnon.dk/how-does-the-nvidia-gtx-1080-stack-up-agains...).
They actually work 15% faster on average compared to the Tesla K series (Remember it's a 5 year old card), but they just stop working after a few months, or return inconsistent results for some operations.
Now, we're not doing graphics with them. We have a GPU database called SQream DB - and we depend on the results to be correct. In the end, they didn't make a lot of sense for us to deploy in a production environment, so back to the Tesla series we went.