I have a 7900 XT, and all my attempts at getting ROCm to work on different Linuxes seem to have failed at different points. Does anyone have a pointer to a clear explanation on how to get it to work?
I have a 7900 XT, and all my attempts at getting ROCm to work on different Linuxes seem to have failed at different points. Does anyone have a pointer to a clear explanation on how to get it to work?
Debian formed a ROCm Team a while ago and we are working hard on getting this to work out of the box, as in: apt-get install and you're done.
For various reasons, our initial focus was on RDNA2 and CDNA2 and earlier, mainly because we are building a CI network in which we test all of our packages (and their reverse dependencies) on actual cards, and we had to bootstrap the network infrastructure first. CI is central to the Debian ecosystem but all our official infrastructure cannot deal with the requirements for specific hardware attached to a machine -- that was never a factor, until now.
AMD has generously supported us with hardware donations, among which are RDNA3 cards. These are already in our possession, and it's just a question of person-time until these are integrated into our CI. We are already spec'ing the systems, though.
As with all other Debian development, we target the 'unstable' distribution for packaging work (similar to git HEAD). Packages eventually migrate from 'unstable' to the 'testing' distribution once they pass certain checks. And when Debian is ready for a new release, 'testing' will basically be tagged as '13 (trixie)'.
Ubuntu periodically syncs their packages from Debian 'unstable', so eventually they will have they same packages that Debian has.
Either way, we realize that running bleeding-edge distributions is probably not for everyone, so we intend to provide backports for both Debian and Ubuntu. Hence why I said earlier that eventually, it shouldn't matter. With "eventually" being sometime in the next few months.
(Site note: our own PyTorch does not yet ROCm support, we're still working on a few dependencies.)
Can I clarify something. If we install a stable release of Ubuntu now, we wouldn't be able to get PyTorch to work with a 7900 XT/XTX? At least not until you've worked your magic and that's made its way to Ubuntu in a few months?
You could still use a version provided by a third party, for example directly from pytorch.org, or a containerized version provided by ROCm upstream [2].
https://www.phoronix.com/news/Radeon-RX-7900-XT-ROCm-PyTorch
(Apologies I missed your response earlier as this thread suddenly had a lot going on.)
After installing the dependencies using apt go to pytorch's getting started page [1] select Stable/Linux/pip/Python/ROCm and run the command
[0] https://rocm.docs.amd.com/projects/install-on-linux/en/lates...
"Hello world" in torch? On a new install/card/config/even upgrade with a bit of fiddling I can get there (I have it more-or-less down at this point).
Actually using something in a real world use case? Hours, days, or give up. Often the third scenario where I end up going back to CUDA because I just need to get something done. Messing around with different driver versions, docker containers, ROCm versions, obscure environment variables, weird random cherry picks/patches and hacks from spelunking a bunch of GH repos/discussions/PRs/blog posts/etc. Feels like it never ends.
I talk about it quite a bit and I keep buying AMD hardware and putting time into this because I'm rooting for AMD. However, consider the CEO of Nvidia saying 30% of their cost/spend is on software. He gets it.
I'm pretty convinced at this point that AMD just doesn't have software in their DNA, with even die-hard "Team Red" Windows gaming users frequently complaining about the quality of drivers. Drivers. For gaming. On Windows. Enter >80% Nvidia market share in that market and >90% market share in ML.
They just don't get it and I'm pretty sure they never will. I expect Intel, Apple, and even new entrants in the field will be the ones to offer viable alternatives to the Nvidia/CUDA monopoly.
Not contradicting you about the other areas, although I do have some hope because things (well... mostly hardware) have already improved in the GPU department under Lisa Su.
Of course with the AI goldrush they're paying a lot of lip-service to ROCm (pumping stock price?) but a couple of examples:
1) They just added official support to ROCm for the top-end consumer card they released over a year ago. Meanwhile in CUDA land Nvidia back-ported support to CUDA 11.8 well before Ada Lovelace/Hopper even launched. Worked with everything out there day of launch. Even newest CUDA 12 and their drivers support anything with Nvidia stamped on it from the past 6-7 years. Literally everything.
2) Just check their repos, Docker Hub, etc. It's embarrassing for them. They used Python 3.10 for ROCm 5.7 because its basically the bare-minimum/standard for any broader AI, ML, etc project. Oh neat, they released ROCm 6 to support their year old ~ $1000 card! Python 3.9? Ubuntu 20.04? What the hell? Pytorch "hello world" usable but almost nothing else even runs.
The further you dig the more clownish it gets.
If I were on an AMD hardware team I would be (mentally) screaming at the software side of the house daily. I can't imagine how frustrating it must be to design and deliver such capable hardware only to have the software render it practically unusable.
Source? According to this https://www.custompc.com/amd-fsr-2-vs-nvidia-dlss-2-comparis... it's only a bit worse:
"However, the results were closer than this might suggest, with both upscaling modes often doing a decent job and delivering very similar levels of image quality. It was just the final 5% where DLSS 2 was that bit better."
> And the end result? Nvidia’s DLSS 2 comprehensively beat FSR 2, coming out on top in almost every comparison and generally exhibiting less ghosting and other visual artefacts while preserving more detail.
If pytorch, keep in mind that most of the pytorch+rocm version is self contained, and you don't need to install every roc* package on your system for it to work. This confused me in the beginning at least.
I totally ignore that recommendation and sometimes have a happily working system and sometimes it's a wreck. In particular, the upstream Linux driver (what you have anyway) and the dkms built one are sometimes rather divergent. Llvm in rocm has different behaviour to the upstream llvm too.
I'm cautiously optimistic that Debian is packaging will end up as turn key just-works. Part of that is persuading AND developers that breaking ABI has consequences, which it kind of doesn't on the install all of rocm together model.
Right now, I use Debian with the same kernel version that the ROCm recommended Ubuntu ships with and build the dkms driver source intended for Ubuntu on it. That be the wrong thing to do but also totally solid in ymmv fashion.
My anecdotal evidence with an AMD GPU is there are serious problems trying to use a graphics card to run graphics and throwing a few compute tasks at it - because X11 will reliably crash and the kernel graphics driver enters some sort of corrupt state. I think that is a use case that was beyond their design considerations to start with.
There is a hint that servers are targeted because they supported linux first with ROCm, and had a very tight support matrix of what OS they expected people to use. Unfortunately, this strategy generates a lot of bad publicity and I'd guess has trouble gaining traction because people can't experiment cheaply with AMD hardware. They're progressing though and even with a dose of cynicism about their wherewithal there is a lot in the pipe that has the potential to take on CUDA. Even APUs are promising.
Honestly, it just stinks of incompetence, or if one were otherwise inclined in beliefs, disgust for their customers.
If AMD had been hiring japanese to make this, they would have had to close up shop as they would have all had to perform seppuku.
It "works", but it's an obvious second class citizen, and the developers don't seem to take bug reports / PRs seriously.
I have the same card, the above works with everything I threw at it so far. Haven't even installed amdgpu-pro drivers btw, only have radv (that steam installed by default).
Edit: Example garbage image: http://oirase.annexia.org/tmp/stablediffusion/sunflower-2.pn...
https://www.phoronix.com/news/Radeon-RX-7900-XT-ROCm-PyTorch