Benchmarking Modern GPUs for Maximum Cloud Cost Efficiency in Deep Learning
minimaxir.com
minimaxir.com
Can someone who is knowledgeable about docker (and container performance in general) comment on how using a docker instance impacts benchmark performance?
I understand containers are not virtual machines, but my preference would be for a benchmark run on bare metal without containerization involved. In particular, it’s not clear to me whether a container can take advantage of the same CPU and GPU instruction optimizations that a binary or compiled-from-source version can. For example, I’m curious if containers package redundant versions of software for convenience that are underoptimized when compared to versions packaged with the operating system. Is that a realistic concern?
In other words, is there generally a meaningful downside to benchmarking with containers, and if so what is it? And equally importantly, are containers the way most people use deep learning frameworks? I have never installed Tensorflow or Keras using a container.
Since the container is running on the bare metal, with the same kernel and everything, just isolated, it absolutely can take advantage of these things. It'll require about the same amount of work to do the compilation for using any special instructions for performance as it would if it's on bare metal.
> For example, I’m curious if containers package redundant versions of software for convenience that are underoptimized when compared to versions packaged with the operating system. Is that realistic concern?
Maybe. It'll mean that there's extra data on disk and maybe in ram from loading other shared libraries which will mean some overhead, but unless there's some specific set of instructions or optimizations that would make a difference for the workload it should be basically the same.
I think this was the parent's worry. Most HPC deployments have versions of libc, libm, etc. compiled with tuning flags specific to the microarchitecture of the system, to allow workloads deployed on them to squeeze as much performance as possible out of the system. I would worry that the versions of libraries in the container are just the generic packages that ship with the OS, compiled with -march=i686 or whatever the modern equivalent is, where they don't try to take advantage of e.g. AVX512 instructions if they're available.
It can be a bit tricky to optimise this though - some deep learning users do things like resizing images on the fly, or on the fly vectorising / scaling of input data in these cases the CPU can be the bottleneck and the optimisations above can help - but this is largely a symptom of a poor pipeline configuration in newbie shops, most serious places have optimised this away and the GPU is the bottleneck
That was essentially my worry, yes. Thanks for clarifying.
The processor won’t (and AFAIK can’t) report different feature sets for processes running in a Docker container, and most heavy optimized numerical libraries will choose to use things like AVX at runtime, not compile time.
We also offer a suite of tools that makes setting up cloud AI pipelines a bit easier to set up and manage.
If this is of interest to anyone here's a $5 promo code to try us out : HNGPU5
full disclosure: I'm one of the co-founders :)
It looks like preemptable GPUs are exactly half the price of normal GPUs (for both K80s and P100s; $0.73/hr and $0.22/hr respectively), so they're about double the cost efficiency (when factoring in the cost of the base preemptable instance), which would put them squarely ahead of CPUs in all cases. (and since the CPU instances used here were also preemptable, it's apples-to-apples)
Spoiler Alert: It is a game changer.
Even if Volta has the speed advantages touted, I doubt P3s will be as cost-effective as a K80.
2. Compiling Tensorflow from source on CPUs is a bit of a hassle but I have seen nice performance gains (10-20%) for LSTM tasks. I bet you would get even higher gains for CNNs since they're more parallelizable. (Note: I've never gotten the latest TF to work with Intel MKL).
3. I haven't fully tested this myself, but with the P100s you also have full support for half precision floats, which supposedly offer a huge speedup.
4. Also would have liked to see benchmarks of other frameworks like PyTorch, etc. I haven't used them myself but everything I've heard indicates that Tensorflow is often slower.
I'm interested to see the utilization of the underlying GPU devices when you run the MLP or CNN benchmarks (monitored with `nvidia-smi`) — the speed-up factor between the different benchmarks don't seem to be inline with the speed-up factor shown in cuDNN link[1] between a K80 and P100. I'm wondering if the P100 device is under-utilized when used with TensorFlow or CNTK.
Separately it seems with ML price/performance is sometimes less important than how much money do you have to spend?
For example with other development work, I’d probably never buy an 18 core cpu because for most projects it wouldn’t speed up my iterative work much.
However ML is more like VFX work, where it’s common that nothing may be too fast. In other words I would be more often willing ignore price/perf, and for example pay 100% more for only a 50% perf gain, if it were within my means to do so.
https://www.ovh.com/ca/en/dedicated-servers/gpu/1801gpu06.xm...
But the best deal with regards to GPUs is to buy your own and put it in your office.
Of course, if you need to use it _all_ the time (or need to heat the space). The whole point of cloud computing is to share those resources with others when you don't need them.
Am I reading the graph wrong, or is this statement not true?
Looks like the P100 is running the code in around half the time as the K80 and the cost is only 150% that of the K80.
.5::1.5 == 1::3 > 1::1