Benchmarking TensorFlow on Nvidia GeForce RTX 3090
evolution.ai
evolution.ai
Edit: It's been edited, thx Evolution :) (or I totally glossed over it the first time around... but I don't think so)
A "Higher is better" might still be interesting although redundant.
It is also not clear what batch sizes are being used for any of the tests. If you switch to FP16 training, you must increase the batch size to properly utilize the Tensor Cores.
If you compare these cards at FP16 performance on large language models (think GPT-style with large model dimension), I am confident you will see Titan RTX outperform the 3090. The former has 130 TF/s of FP16.32 tensor core performance while the latter has only 70 TF/s.
Link: https://www.nvidia.com/content/dam/en-zz/Solutions/geforce/a...
So if you're someone who does their work mainly in FP32, you will see improved performance with the 3090. On the other hand, if you are an FP16 speed demon who needs to train GPT-3 over the weekend, stick with your Titans :)
But on 3090, I don't think the speedup will be 5x, it should be closer to like 2x. The 3090 has 35.6 TF/s at TF32 and the Titan RTX has 16.3 TF/s at FP32. Once again I think there is handicapping going on for 3090.
Also, why FP32? CNNs are some of the most robust models to train in FP16 (much easier than language models) so you could get yourself a quick XXX speedup and 2x memory savings by switching over.
(btw not intending to be accusatory or anything, I just think FP16 training deserves a lot more adoption that it currently seems to have :)
In games the 3090 only gives a 15% performance bump relative to the 3080. If that pattern holds for machine learning tasks there is probably a scenario where it makes sense to buy two 3080s rather than one 3090.
If you are vram constrained then obviously the 3090 the way to go.
Could you kindly advise what kind of computer would make sense to purchase to begin learning about ML? I was assuming I'd get a 3080. Should I get a case that could potentially house 2 x 3080's? Does the case require any special cooling considerations, or just whatever will fit the cards? What CPU would you get?
And if you ARE actually running it that hard, you'd better budget for fairly frequent replacement cards.
And no, cards don't just keel over in a few months at 100%. Crypto miners ran that experiment. A typical card has years of 100% in it.
- Someone else is paying
- You expect to dabble
- You need burst capability
Buy if:
- Cost sensitive & capable
Are modern cards really so fragile you can expect them to die off under heavy use even if properly cooled and not overclocked?
Just buy a PC that you like for gaming(with an Nvidia gpu) and don't worry about ML yet - it's incredibly unlikely that you can pick something that would limit you in any way. Small datasets will run on anything, large datasets will take hours to process no matter what you run them on. It's not a "limit".
http://timdettmers.com/2018/12/16/deep-learning-hardware-gui...
make sure your motherboard and processor support whatever the newest version of PCIe is -- a major factor with deep learning is bandwidth moving data on/off the GPU.
AMD GPUs can theoretically be used for machine learning, but right now software support is lacking -- you will spent more time configuring and installing than learning. (AMD CPUs are fine though.)
it doesn't really matter that much though -- any gaming PC with a new-ish NVidia card can be used to do quite a bit of interesting ML.
Nvidia came to dominate the market at a time when AMD wasn't making particularly competitive GPUs, but that isn't really the case anymore. For anything not so expensive that nobody is really going to buy it anyway, the current and expected (in less than a month) AMD GPUs are competitive on performance.
The result is that a lot of large customers, who see value in not being locked into a single supplier, are going to be pushing for frameworks that work across multiple vendors. And then you could plausibly be wasting your time learning Nvidia-specific technology which is about to become disfavored. So you might want to wait and see.
Twice bitten... once shy? In any case, I'm going to let someone else be the guinea pig this time.
so the question is just -- when will it be very simple to install these packages for AMD GPUs, with enough mathematical operations implemented and optimized to let you do the things you want to do.
right now things sort of work, but it's definitely in a bleeding edge early adopter state. it's seemed like AMD is on the cusp of catching up for a couple years now, but it's taken longer than I expected.
True, but even once TF/PyTorch support AMD well it's highly possible that an unanticipated CUDA dependency will pop up in one's computational journey. NVidia subsidized CUDA seminars for a decade and now it's all over the place, both in the flagship frameworks and in the nooks and crannies.
GPUs are only really required in ML if you want to do deep neural network stuff. You can do plenty in CPU on reasonable data sets using any modern laptop.
If you’re just wanting to learn machine learning you don’t need anything particularly special. I think you would be happy with GTX 1070. There is also the cloud computing route where you basically rent the gpu from AWS. That will be initially more cost effective than buying your own hardware.
One thing to keep in mind if you do go with the 3080 is the power consumption. Ampere cards are going to be much more power hungry than previous generations, and you will need to budget about 320W just to the graphics card. The recommended power supply for the 3080 is 850W.
This has to do with transients, the 3000 series cards have some massive transients that can easily trip OCP protection on powersupplies not designed for that kind of transient. A 700W powersupply is able to handle those transients much better than a 500W PSU is.
Of course if money is an issue, you are well off with only a single 3080 :)
There is no point in buying the 20xx series anymore. The 30xx are twice as good for the same money (if you can get one).
You can get very far on any laptop before hardware becomes the main blocker. And before building an ML machine, there are cloud compute options available for far cheaper.
- this is a sign that you are most likely doing it wrong. Yes, some operations are inherently bandwidth bound, but most important ones such as larger matrix multiplies (transformers) and convolutions are compute bound.
https://timdettmers.com/2020/09/07/which-gpu-for-deep-learni...
HN Discussion:
Edit: clarified that I am referring to slower relative performance
For “tensor ops” in GeForce cards FP16 with FP32 accumulate is done at half rate so you don’t get double the performance which you do get in Quadro and Titan cards using the same die.
https://www.techpowerup.com/gpu-specs/geforce-rtx-3090.c3622
But does model get quality hit: need to train for more steps before converging to the similar performance and have more parameters?
FM16 obviously contains less information than FP32.
+ tf-nightly and other python libraries installed through pipenv
It helps keeping to Ubuntu LTS versions though, that's what they support best.
A couple of months ago, I removed all the references in apt sources, and followed the newer instructions (several times to get the right driver/cuda/tensorflow match) and my reboots are great, and only one GPU lock up so far (probably due to overheating - I've had to replace a couple of components flag as failed due to the heatwave in summer)
Jupyter hub is just great, I'd like to implement better diagnostics though ... have yet to find a good tutorial for that as yet.