Which GPUs to Get for Deep Learning
timdettmers.wordpress.com
timdettmers.wordpress.com
I work with data sets > 250GB: GTX Titan eBay
I have no money: 3GB GTX 580
I do Kaggle: GTX 980 or GTX 580
I am a researcher: 1-4x GTX 980
I am a researcher with data sets > 250GB: 1-4x GTX Titan eBay
I never used deep learning before: 3GB GTX 580 eBay
I started deep learning and I am serious about it: Start with one GTX 580 and buy more GTX 580s as you feel the need for them; buy new Volta GPUs in 2016 Q2/Q3
I want to build a GPU cluster: This is really complicated, I will write some advice about this soon, but you can get some ideas here"
While I was really looking forward to Intel shipping a competitive processor to NVIDIA's GPUs, reality disappointed. Intel put 2 engineers on porting a piece of code I wrote for a year. That code is roughly 20x faster on a Maxwell class GPU than on a Haswell class CPU (all cores firing). 2 man years later, instead of the Xeon Phi performing on par with a C1060/GTX 285 as it initially did, it managed to accelerate the CPU code by ~10%. In that time, NVIDIA shipped GK104, GK110 and then GM204 accelerating my code by a factor of 2 relative to GK104. When I tried to help one of the Intel engineers catch up to at least Fermi class GPUs, he got mad at my algorithmic choices and stormed out of the room.
Intel IMO will resort to dumping these things at rock bottom prices into data centers in order to spoof just enough low information government wage HPC administrators into believing these things are good for anything more than playing Jenga with them. And that will keep NVIDIA from achieving total HPC dominance in the short term.
No idea what the long-term will bring. Intel ought to be able to deliver a compelling part one day, but I suspect that will be a CPU with wider SIMD and more cores.
The general consensus I've seen is to just get an NVIDIA card if you're serious about working with deep neural nets on the GPU.
One thing that did surprise me was that there was no mention of using EC2 GPU spot instances for getting your feet wet. If you don't have access to a GPU with CUDA support you can get a spot instance for about $0.07 an hour to at least test out that you have your GPU code configured correctly (and you will see some performance gains). There are even a couple of AMIs out there with Torch7 and Theano already installed.
AWS is great if you want to use a single or two separate GPUs. However, you cannot use them for multi-GPU computation as the virtualization cripples the PCIe bandwidth; there are rather complicated hacks that improve the bandwidth, but it is still bad. Everything beyond two GPUs will not work on AWS because their interconnect is way to slow.
Where are the citation for this? I have heard about the disappointing benchmarks when the 600 serie launched, but this is uncommon as far as I know in the GPU marked to have sub part drivers on launch. Plus if you are following a tic-toc strategy like Intel does, you should rarely go for the tic iteration in my opinion. So is this just a claim from the author or is there something to back this statement up?
Completely my fault, but still disappointing
So the 0.5GB region is still faster than PCIe IIRC. Not ideal of course, but still better than a bunch of PCIe transfers.