Would you still do it without hesitation if the money was coming of your own pocket?
Would you still do it without hesitation if the money was coming of your own pocket?
AMD GPUs have near zero support in major ML frameworks. There are some things coming out with Rocm and other niche things, but most people in ML already have enough work dealing with model and framework problems that using experimental AMD support is probably a no go.
Hell, if AMD had a card with 8GB ram more than nvidia, and for 500$ cheaper, I would still go with nvidia. Everyone wish AMD would step their game up w.r.t ML workloads but it's just not happening (yet), Nvidia has a complete monopoly there.
So AMD got a lot of heat for not supporting Navi (RDNA) with ROCm, but it seems that they are weeding out the things keeping people from running it ( https://github.com/ROCmSoftwarePlatform/pytorch/issues/718 and the links in that look like gfx10 is almost there for rocBlas and MIOpen). We'll see what ROCm 3.9 will bring and what the state of big navi is.
Why would I get a Radeon VII when used nvidia cards for machine learning are extremely cheap, and then I don't have to worry about experimental stuff breaking one day before deadline lol
_If_ AMD has made a sufficiently powerful GPU, that will add a lot of incentive to ML frameworks to support it. But it's going to have to be a big difference, I imagine.
Given how active AMD is in open source work, I'm a little surprised they haven't been throwing developers at ML frameworks.
* It seems there was not a single commit in the past ~6 months which by itself is already a deal breaker.
* Documentation is lackluster
* You need to use Keras, I use PyTorch. This is not a deal breaker, but a significant annoyance.
* Major features are still lacking. E.g. No support for quantization (afaik), which for me is fundamental.
* Most importantly there seem to be no major community around it.
It feels a bit bad to say that, because clearly a lot of work went into this project, and some people have to start adopting it to drive the momentum, but from an egoistical point of view, I just don't have the courage to deal with all the mess that comes with introducing an experimental layer in my workflow. Especially in ML where stuff can still appear to "work" (as in, not crashing) despite major bugs in the underlying code, leading to days or week of lost work before realizing where the issue is.
Check out keras-helper.. I made it to switch between various backend implementations which are non NVidia specific.
Pytorch may eventually need porting, but for now I don't need it. I've been trying out Coriander and DeepCL now but I decided to stick to Keras, which seems to be a decent compromise. Not using 2.4.x though, do not need it.
OpenCL based backends are cutting it for me, running production workloads without needing to install CUDA/ROCm is the best way to go.
Spending a week/year working around ROCm would already cost you 5k$ plus the opportunity cost. For a whole team that’s a money sink.
The analogy is that everyone should get NVIDIA Ampere units (non consumer) units worth $30k because it's fast and you'd rather be spending less time in a lab with millions of dollars in funding. insert don't be poor T Shirt reference
PlaidML is not ROCm. Nobody needs ROCm, what people need is just linear algebra well implemented with OpenCL primitives. That's what PlaidML is. And it works quite well, even on those integrated Intel GPUs on most laptops.
Have you also looked at DirectML and WSL2? They seem to be running tensorflow quite well too. Those things may be the key to bringing these in adoption outside the well paid class of data scientists you came up with.
The catch is that ML software stacks have had hundreds if not thousands of man-years of effort put into things like cuDNN, CUDA operator implementations, and Nvidia-specific system code (eg. for distributed training). Many formidable competitors like Google TPU have emerged, but Nvidia is currently holding onto its leadership position for now because the wide support and polish is just not there for any of the competitors yet.
That is incorrect - I have been running Tensorflow on RX5*0 cards for close to 2 years now. I even transitioned to TF2 with no problem. Granted, I have to be extra careful about kernel versions, and upgrading kernels is a delicate dance involving AMD driver modules and ROCm & rocm-tensorflow. My setup is certainly finicky, but to say AMD GPUs have near zero support is false.
gimme 48 GB gimme 128
Now on the other side of the world, where a 3090 is several times your rent, you really need to think thrice about buying one.
edit: Also CUDA is just too important to switch to AMD.