What Is AMD ROCm?
threedots.ovh
threedots.ovh
AMD had no support. Card maker said this didn't fall under warranty. I got burned over and over.
I bought NVidia. It just worked.
I'm working on a potentially major piece of infrastructure, and AMD is accumulating debt. If it worked out-of-the-gate, I imagine we would have kept support. Within 6 more months, we'll be NVidia-specific. AMD will be that much further in the hole for support.
I'd love for ROCm to win, since I think open is critical here. On the other hand, I can't imagine it will. AMD would need to run this as a loss leader for a while, and engineer this at a level to get this competitive with NVidia.
A half-baked product like ROCm seems like a money hole for everyone involved. Customers get burned, and I can't imagine AMD comes out positive.
In the meantime, NVidia is minting gold here.
Now I'm in "twice bitten, once shy" mode with AMD. I hate paying the green tax as much as the next guy and I desperately want to have a second source of professional GPUs, but I'm not going to be the guinea pig. Not again. Not for the 3rd time. I want to see someone else successfully using AMD cards for common ML workflows and for blender before I even consider risking it again.
Their problem has been that as little as five years ago they were a dying company. They didn't have the resources to do this right.
That's no longer the case, but once you have the money there is still a lag between then and when the release funded by that money comes out. And even then they're fighting an uphill battle against the perceptions created during their dark age.
Probably the biggest thing they have working for them is Nvidia's behavior. Proprietary everything and single vendor lock in makes everybody chafe, so as soon as they can produce something usable, everyone will want to use it.
I think that will be an increasingly hard bar to clear, as software becomes coupled to CUDA, though. AMD will be chasing a losing race. They won once with Intel, but this one feels harder....
And even if AMD's offering was not an absolute dumpster fire, Google, Microsoft and Amazon all have their own accelerator that are maturing and will be more cost effective on the long run.
It it is the usual Khronos defines the base stuff and hopes for the best regarding their partners.
I want to use it for compute on something like a rx 6800 and to my knowledge can't
Not really what you expect from quality engineering. At the end of the day these kind of companies don't understand the value of development and engineering clients as customers.
It's unfortunate really.
EDIT: here's an example:
ROCmSupport commented on Feb 22 •
Hi @powderluv
Thanks for reaching us. I can not comment on RDNA2 support right now.
We are working on adding a few more new hardware into ROCm environment.
Please stay tuned via our documentation.
Thank you.
@ROCmSupport ROCmSupport closed this on Feb 22
https://github.com/RadeonOpenCompute/ROCm/issues/1390#issuec...I guess on the plus side they at least have a more open driver than NVidia (AFAIK nouveau doesn't get any support from them, at least AMD tries to maintain their open source driver on some level.)
And yet, every time I've tried an ATI/AMD Card, the driver experience even in windows has been pretty off-putting, and while I suppose we are finally at a point where one is less likely to be impacted by their issues with 768p overscan on TVs, I wonder what zany quirk they'll come up with next.
I think you somehow mistyped "AMD has open source drivers of absofuckinglutely excellent quality supporting hardware of the last ten years or so". No really, they are great. For graphics, that is.
(I work for AMD on ROCm. All opinions are my own.)
I doubt they have the funding to meaningfully impact the overall hardware and software support matrix but if they could just make the GitHub repo feel less like I’m back in my days working at a call centre raising tickets to a second level support team in a foreign country who’s only business KPI was tickets closed per day.
I think it's worth noting, though, it's not always as bad as the example in the sibling comment. The RadeonOpenCompute/ROCm repo catches a lot of questions about big features and the future direction of the project. Those are particularly difficult to answer as an engineer. As much as I'd like to, I can't make a product announcement in a GitHub issue.
If you have a specific technical problem and you open an issue on the repo for the corresponding component, you'll probably have a better experience. Some teams are more responsive than others, but that will at least maximize your odds of successful resolution.
With CUDA you simply target a specific CUDA version and there is full forward and backwards compatibility on any hardware that supports that version.
Once the developers are familiar with CUDA, what are the chances you'd choose ROCm for deployment? Yeah, not great.
Consumer cards' ROCm support is strategic.
Given that compute is important, and many people use their GPUs for compute of some sort, I simply can not understand how that part is so poorly executed on the part of AMD that, in terms of actual application, you might call it entirely absent.
I mean, this is a company that produces compute cards, which supposedly someone in the world must buy and use... but who? Why? I have never seen anyone, and for good reason. And it seems like AMD just... doesn't care?
Like, it's just not a part of their organizational strategy... Compute is on the powerpoint slides, but... no one (can) use it?
It's been going on for years now and I don't get it.
They announced the MI200 GPU with 128GB of memory, two supercomputers (Setonix, Frontier) are suppose to include them but both will only be launched next year https://www.tomshardware.com/news/setonix-supercomputer-mi20...
Cuda on linux with ml/gpu workloads is still kind of a hotmess and i'd say we're far from finding a winner like some suggest here.
It's gotten better... but still far easier to treat it like a mess and start fresh with any install
Almost everyone on Linux will have experience of breaking their drivers at some point, and installing another alternate set is a big risk.
It seems silly to not have OpenCL and HIP access without having to use this alternate stack.
So even if I wanted, I can't. Sincerely, I'm fed up with AMD's attitude towards compute.
ROCm 4.5 is the _last_ version to support the Vega10 ASIC (MI25, Vega56, Vega64).
https://github.com/RadeonOpenCompute/ROCm/#amd-instinct-mi25...
The next ROCm release after 4.5 is sometime in Q1 next year. So it's on planned death really soon.
It is transitioning to _that_ comical AMD "enabled in the codebase but not tested and not supported" state, rotting slowly like Polaris support did.
I was concerned about that as well, but I don't personally know anyone who owns a gfx900 card. I'm a little unclear on what impact it will have on the community.
The MI25 wasn't targeted at the general public, but it wasn't hard to buy one. And the customer products using that same die, Vega 56 and 64, were sold quite a bit.
What affected the community severely might be the combination of both GFX8 and gfx900 going away, leaving only MI50 (Vega20, also used in Radeon VII consumer variant) and no support (only unofficial enablement) for the later customer products. Because those went GFX10/Navi.
[0] https://en.wikipedia.org/wiki/GPUOpen#Radeon_Open_Compute_(R...
The old expansion still shows up in a few places where it's difficult to remove.
[1]: https://github.com/ROCmSoftwarePlatform/rocBLAS#rocblas