Which GPUs to get for deep learning
timdettmers.com
timdettmers.com
Unfortunately the performance charts are completely devoid from reality, and in particular the discussion on tensorcores may be true from an instruction count perspective but does not reflect any third-party benchmark I’ve seen. For example: https://lambdalabs.com/blog/2080-ti-deep-learning-benchmarks.... Nvidia has a history of straight-up-lying about tensorcore and other benchmarks (for example, see this thread from right after Nvidia announced an 8x improvement in speed on imagenet in tensorflow for V100: https://github.com/tensorflow/benchmarks/issues/77)
In general, fp16 is only 30-40% faster than fp32, and occasionally 2x in really optimal conditions.
The performance numbers posted here appear to almost exactly reflect the LambdaLabs numbers.
Lambda Labs: the RTX 2080 Ti is 96% as fast as Titan V, 73% as fast as Tesla V100 (32 GB)
timdettmers: RTX 2080 Ti normalized to 1, Titan V looks about 1.1 to 1.2, V100 is just below 1.5
> for example, see this thread from right after Nvidia announced an 8x improvement in speed on imagenet in tensorflow for V100
Well a non-NVidia person managed to get it up to just above 4x improvement without "using unreleased libraries from NVIDIA". From the same thread: https://github.com/tensorflow/benchmarks/issues/77#issuecomm...
In my experience NVidia benchmark numbers in deep learning are rarely lies - they are highly optimised, in optimal conditions and rarely achievable in the real world. About what you'd expect from a vendor benchmark.
> In my experience NVidia benchmark numbers in deep learning are rarely lies - they are highly optimised, in optimal conditions and rarely achievable in the real world.
Right, but Nvidia claimed 1360 images/sec for resnet-50 on imagenet. To my knowledge this still hasn’t been realized by a third party. It also isn’t a 4x improvement for fp32 vs fp16, that’s comparing to previous generation. Improvement is more like 1.5x: https://lambdalabs.com/blog/best-gpu-tensorflow-2080-ti-vs-v...
Even in very simple synthetic benchmarks the speed up is only 2x: https://github.com/tensorflow/benchmarks/issues/77#issuecomm...
I have not seen any benchmarks showing an 8x speedup. Have you? If not -> Nvidia lied.
On https://images.nvidia.com/content/technologies/volta/pdf/vol... they claim 1,525 images/sec (!)
Dell hit 5,243 images/sec with one of their 4 V100s servers, which comes to 1,310 images/sec per V100. I find it very believable that NVidia would get ~200 images/sec more, since Dell jumped 50% with a change in their CPU/GPU connection topology.
See https://www.dell.com/support/article/en-au/sln317397/deep-le...
The thing I've seen is very limited (NVIDIA GPUs offer up to 8x more half precision arithmetic throughput when compared to single-precision, thus speeding up math-limited layers.[1]) which is probably true.
The problem with performance improvements is the diminishing returns part of Amdahl Law: an 8x improvement in math performance will just mean the math part becomes less important in terms of absolute performance.
In any case, I've found NVidia's claims in the machine learning area to be pretty good. Like most claims you have to read carefully to see exactly what the claim is, but that's not uncommon with performance claims.
[1] https://docs.nvidia.com/deeplearning/performance/mixed-preci...
[2] https://en.wikipedia.org/wiki/Amdahl%27s_law#Relation_to_the...
Right but I’ve benchmarked the best case scenario, ie a large GEMM call in C++, and still not seen anywhere close to 8x. I’ve never seen a code example, no matter how limited, showing a 8x speed up.
See cuBLAS mixed-precision GEMM.
It really depends on workload. For ImageNet and resnet type architecture its not unusual to get 3X speed up. It also depends on if you do full fp16, leave BNs alone etc. This is a lot because instead of training for 6 days, you now training for 2 days.
Source on this? I've done a good bit of CV benchmarking work and I don't recall anything like a 3x boost. 30-40% improvement is much more in line with what I remember.
One thing that I am quite sure of for the A100 is its transformer performance. It turns out, large transformers are so strongly bottlenecked by memory bandwidth that you can just use memory bandwidth alone to measure performance — even across GPU architectures. The error between Volta and Turning with a pure bandwidth model is less than 5%. The NVIDIA transformer A100 benchmark data shows similar scaling. So I am pretty confident on the transformer numbers.
The computer vision numbers are more dependent on the network and it is difficult to generalize across all CNNs. For example, group convolution or depth-wise separable convolution based CNNs do not scale well with better GPUs and speedups will be small (1.2 - 1.5x) whereas some other networks like ResNet get pretty straightforward improvements (1.6x-1.7x). So CNN values are less straightforward because there is more diversity between CNNs compared to transformers.
This strategy can keep you going for a couple of years. With models becoming as big as they are, I doubt how much SOTA an RTX 3070 is going to do in 2-3 years. (actually none of these cards come close to GPT-3). By that time you can pick up a second hand RTX 30x and still get the latest offerings in the cloud.
Buying second hand GPU comes with a bit of a risk by the way, someone suggested only buying if the price is less than half of the original price (can't remember the link).
> Do not buy GTX 16s series cards. These cards do not have tensor cores and, as such, provide relatively poor deep learning performance. I would choose a used RTX 2070 / RTX 2060 / RTX 2060 Super any day over a GTX 16s series card
...a few paragraphs later...
> If that is too expensive, a used GTX 980 Ti (6GB $150) or a used GTX 1650 Super ($190).
If you are on a tight budget, his advice is to pick in this order:
> I have little money: Buy used cards. Hierarchy: RTX 2070 ($400), RTX 2060 ($300), GTX 1070 ($220), GTX 1070 Ti ($230), GTX 1650 Super ($190), GTX 980 Ti (6GB $150).
I figured that most people that start with deep learning might also lack cloud computing skills. Learning one thing at a time is easier and as such, just sticking a GPU into your desktop and focus on deep learning software / programing might yield a better experience.
I might update my blog post in the future with this detail.
Fun fact: I've saved your post as PDF for offline reading and it clocked at 649(!) pages at the time of saving (32 pages for the post per se and the rest for the blog post comments). Combining that with feedback here at HN, it is clear that there is quite a lot of interest in the topic ...
I should also a bit more in general about cloud computing, it seems some people agree that the post ran a bit short on that. At some point I just wanted to be done with it though — editing 10k word blog posts is not so much fun anymore!
I can certainly understand you being hesitant to add more content to an already sizeable post. Perhaps, several small paragraphs on important relevant aspects might still be worth considering (take it with a grain of salt, since I haven't actually read your post in detail, including cloud-related parts). Anyway, thank you very much, again, for your time and effort. Keep it up!
I've won multiple silver Kaggle medals on a 1070. It's true that more power would be helpful, but I feel it's lack of technique (and time!) rather than compute that has held me back from gold medals.
I've thought about it a lot, and talked to lots of really good Kagglers about it. Most of them use multiple machines, rather than having maximum performance in a single machine.
This lets them run multiple completely different experiements at once, and then put extra compute onto the ones that seem good.
That is a big difference to me, when I have to experiment sequentially. The parallelism is more important than absolute speed a lot of the time.
I am also interested in whether PCIe4 can help unified memory for larger models. Guess have to wait for RTX 3090 actual release.
What are you talking about? Are you writing custom CUDA code to run those large models? Because there's zero support for unified memory in any of the existing DL frameworks.
When I tested the DGX2, 16-card bandwidth was finally not an issue (I was extremely happy with the NVSwitch)... but I hit a CPU bottleneck going from [64vcpu + 8 V100] to [96vcpu + 16 V100] (even with tricks to reduce CUDA CPU load).
I’m excited to sometime try my workload on 2x NVLinked 3090s with a 64-core AMD CPU, optimized PCIe v4 NVMe data path, additional caching, and some hand tuning of the training code. It’s possible this will be competitive with an 8x V100 google cloud instance based on some of my scaling pain graphs.
I’d consider 4x 3090s- but between only having a single NVLink port, and having to figure out how to even fit four 3x cards in a case without water cooling, it seems prudent to start with two.
In short: You need lots of RAM.
And stay away from overclocked (founders edition) and from datacenter models due to heat or price problems.
And also cost $$$ which makes it a luxury solution only suitable for the users who absolutely want a thin portable notebook for on the go and a powerful GPU at home for AI/Gaming.
I have a box with four of them which have served me well for a while now, and based on just raw performance it looks like upgrading to two RTX 3080s would exceed the performance of my current system.
I'm wondering if I should rush to sell off the cards on the used market before the prices crash and then use that money to swap over to Ampere.
Then, there's also the question of if there will be an RTX 3080 TI which will blow away the RTX 3080 and be a viable card for the next five years like the 1080 TI.
I'm really uncertain about what to do and wonder the calculus other people have done on this decision.
I've gone with GPU spot instances for my personal experiments. The key is to be able to bring up a machine in a couple minutes so you're ok with always tearing one down. A combination of ansible and some scripts that push code around helped a lot to create a useful environment for experimenting.
It would be great if you could add something about the hardware requirements of the reinforcement learning and video prediction.
Though, you should probably use AWS ML Compute, since they even have Nvidia Ampere 100's, which cost $10,000 each, and it'll probably be more cost effective for heavier workloads.
Hope it will not ... 3090 got 24 and price is not as unreachable as Titan then.
NVidia cannot be blamed for their incompetency.
Why didn't AMD do this as well? I thought AMD was strong on supporting open source.
No, you don't _need_ to use CUDA for GPU programming, you can use OpenCL or Vulkan or probably even PHP instead.
I do, however, _want_ to use the best programming language for the task at hand. If that task is GPU programming, CUDA is the best language I know for that, much better than SyCL, OpenCL, Vulkan / OpenGL + shaders, etc.
If these other technologies would be better, I would use them instead.
Not that this matters because your argument is flawed.
The claim that CUDA is not worth using because it lacks portability, only holds, if there is hardware worth using that's not supported by CUDA.
The only GPUs worth buying for compute are from nvidia and support CUDA, so your claim isn't true.
The only thing you achieve today by not using CUDA is paying a huge price in development quality for portability that you can't use.
The startup cemetery is filled with companies that made this trade-off and picked OpenCL just in case they wanted to use non-nvidia hardware. They were all killed by the velocity of their competitors that were using CUDA to deliver better products that payed the bills.
The only people for which it might make sense to avoid CUDA are "non-professionals" (hobbyist, etc.). If you only want to use OpenCL to "learn OpenCL", then OpenCL is the right choice. But if you want to make money, then CUDA was the right choice 15 years ago and still is the right choice today.
If that makes you angry, direct your anger properly. It isn't NVIDIA's fault that CUDA is really good. It is however, AMD's, Intel's, Apple's, Qualcomm, ARM's... fault that everything else _sucks hard_. Being angry at nvidia for delivering good products is just stupid. Its the other companies fault that they can't seem to be able to get their sh* together when it comes to GPU computing.
That sounds like koolaid marketing to me. AMD GCN was more compute oriented than Nvidia for years and only lately AMD increased focus on gaming with RDNA.
That's a fact: check HPL, MLPerf, Spec, etc. results. MLPerf is the perfect example, were your results are only accepted if they can be verified by others. Where is AMD in there? (nowhere, their products suck for compute).
> AMD GCN was more compute oriented than Nvidia for years
No, the only thing AMD GCN was good for is as a very expensive stove.
AMD GCN had a lot of compute, on paper, and higher numbers than nvidia GPUs of the time. Unfortunately, AMD GCN's memory subsystem sucked, and it was impossible to deliver data fast enough to actually be able to use the compute.
So nvidia's hardware essentially destroyed GCN for any useful practical application.
IIRC, the only application for which GCN's got some use was bitcoin mining, which avoided hitting GCN's issues because it just requires doing a ton of useless work on a tiny amount of memory. Perfect for GCN right? Nope, nvidia's hardware was still better, but sold out, and GCN wasn't horrible at this, so it got some use.
AMD actually fired the architect of GCN over this. Yet this still perfectly summarizes AMD's GPGPU strategy of the last 15 years: higher numbers on paper, that cannot be achieved in practice, and lower that the numbers that nvidia's hardware achieves in practice.
There aren't any.
https://web.archive.org/web/20200907164516/https://timdettme...
This post looks very thorough, and came just in time for me. I'm looking to snag an upgrade from my GTX 970 for a mix of flight sim 2020 and digging into Fast.ai's course part 2.
The 970 has been my big hold-up, right now even simple models take a really long time to work with.
I really dig the overall idea of cloud notebooks. Back when I did fast.ai part 1, I used Paperspace Gradient. It was a pretty good experience, but moving files around was a bit of a hassle. For example, getting the images for the Planet Labs exercise took a round trip of downloading from Kaggle to my computer and re-uploading into Jupyter to do analysis.
Because of all those moving parts, I decided to give a try running things locally. To my surprise, setup was super easy and I was quickly productive! I really dig how customizable a local Jupyter server is, too.
I do use Colab, it's particularly great for collaboration/sharing notebooks, but my past experience has me hooked on the idea of a capable ML machine at home.
Plus: I can pitch it to myself and my spouse as an investment in personal development that happens to be able to game :D