Maybe they should look into their own failures first.
Maybe they should look into their own failures first.
https://www.intel.com/content/www/us/en/developer/tools/onea...
My work involves writing software that runs on many GPU platforms at once. So far we have been going through the Kokkos route, but SYCL is looking pretty good to me recent days. There is some consolidation happening in this space (Codeplay gave up working on their own implementation and merged with Intel). It was pretty easy to setup on my Linux machine for Nvidia card. Documentation is very good and professional, unlike AMD's, which can be frankly horrible at times. And Intel has a good track record with software.
I genuinely believe if someone is going to dethrone CUDA, at this point SYCL (oneAPI) is a far more likely candidate than Rocm/HIP.
This would be the first step. Then, if we want to move away from Cuda into hardware that's as ubiquitous and performant as Nvidia's (or better), someone would need to write an abstraction layer that's more convenient to use than Cuda. I did play a little bit with Cuda and OpenCL, but not enough to hate either.
Any Cuda-compatible software layer only has two options, be a second class CUDA implementation by being compatible with a subset like AMD ROCm and HIP effforts, or be compatible with everything always playing catchup.
The only way is to use middleware that just like in 3D APIs, abstract the actual compute API being used, as man language bindings are doing nowadays.
Implementing programming languages and runtimes is pretty difficult in general. Note that cuda doesn't have the same semantics as c++ despite looking kind of similar. Wherever you differ from expected behaviour people consider it a bug, and implementing based on cuda's docs wouldn't get you the behaviour people expect.
Pretty horrendous task overall. It would be much better for people to stop developing programs that only run on a gnarly proprietary language.
Yet another example on how Intel and AMD failed to take up on CUDA.
SYCL (SYCL-2020 spec) supports multiple backends, including Nvidia's CUDA, AMD's HIP, OpenCL, Intel's Level-zero, and also running on the host CPU. This can either be done with Intel's DPC++ w/ Codeplay's plugins, or using AdaptiveCpp (aka. hipSYCL, aka openSYCL). OpenCL is just another backend.
It is also a very long way from OpenCL C++. The code is a single C++ file, and you don't need to write any special kernel language. The vast majority of SYCL is just C++, so -if you avoid a couple of features- you can use SYCL in library-only form without even any special compiler! This is possible for instance with AdaptiveCpp.
Also some of that work, we have to thank Codeplay for, before their acquistion from Intel.
I'm not an expert here, am I missing something? Saying the x86 industry is motivated to move away from what nvidia provides, intel needs to tick some of these 'better somehow' boxes.
I attribute this to opencl being the common subset a bunch of companies could agree could be implemented. I wrote some code that compiles as cuda, opencl, C++ and openmp, and the entire exercise was repeatedly "what, opencl can't do that either? damn it".
Apple also doesn't care about HPC, and has their own stuff on top of Metal Compute Shaders, just like Microsoft has DirectML.
Unless a random gamer with a random AMD GPU can go to amd.com and download pre-packaged, officially supported tools that work out of the box on their machine and after a few clicks have working GPU-accelerated pytorch (which IMHO isn't the case, but admittedly I haven't tried this year) then their "doubling down" isn't even meeting table stakes.
I predict that access to the newer cards is a more likely scenario. Right now, you can't rent a MI250 or even MI300x, but that is going to change quickly. Azure is going to have them, as well as others (I know this, cause that's what I'm building now).
I'm considering adding ROCm support for some ML-enabled tool - no matter if it's a commercial product or an open source library - the thing I need from AMD is to ensure that the ROCm support I make will work without hassle for these end-users with random old AMD gaming cards (because these are the only users who need the tool to have ROCm support), and if ROCm upstream explicitly drops support for some cards because AMD no longer regularly test it, well, the ML tool developers aren't going to do that testing for them either; that's simply AMD intentionally refusing to do even the bare minimum (do a lot of testing for a wide variety of hardware to fix compatibility issues) that I'd expect to be table stakes for saying that "AMD is doubling down on ROCm".
ROCm is a stack of a whole lot of stuff. I don't see a stack of software being "the whole point".
> the thing I need from AMD is to ensure that the ROCm support I make will work without hassle for these end-users with random old AMD gaming cards (because these are the only users who need the tool to have ROCm support)
From wikipedia:
"ROCm is primarily targeted at discrete professional GPUs"
They are supporting Vega onwards and are clear about the "whole point" of ROCm.
> "ROCm is primarily targeted at discrete professional GPUs"
That's kind of true, and that is a big part of the problem - while AMD has this stance, ROCm won't threaten to replace or even meet CUDA, which has a much broader target; if you and/or AMD want to go in this direction, that's completely fine, that is a valuable niche - but limiting the application to that niche clearly is not "doubling down on ROCm" as a competitor for CUDA, and that disproves the TFA claim by Intel that "the entire industry is motivated to eliminate CUDA", because ROCm isn't even trying to compete with CUDA at the core niches which grant CUDA its staying power unless it goes way beyond merely targeting discrete professional GPUs.
It isn't just old cards though, CUDA is a point of centralization on a single provider during a time when access to that providers higher end cards isn't even available and that is causing people to look elsewhere.
ROCm supports CUDA through the included HIP projects...
https://github.com/ROCm/HIPIFY
The later will regex replace your CUDA methods with HIP methods. If it is as easy as running hipify on your codebase (or just coding to HIP apis), it certainly makes sense to do so.
What they really need is to support the less expensive cards, of which the older cards are a large subset. There are a lot of people who will make contributions and fix bugs if they can actually use the thing. Some CS student at the university has to pay tuition and therefore only has an old RX570, and that isn't going to change in the next couple years, but that kind of student could fix some of the software bugs currently preventing the company from selling more expensive GPUs to large institutions. If the stack supported their hardware.
Support was dropped upstream, but only because AMD no longer regularly test it. The code is still there and downstream distributors (like if you just apt-get install libamdhip64-5 && pip3 install torch) usually flip it enabled again.
https://salsa.debian.org/rocm-team/community/team-project/-/...
There is a scary red square in the table, but in my experience it worked completely fine for Stable Diffusion.
Those kids and hobbyists can't even rent time on high end AMD hardware today. I see that as one piece of the puzzle that I'm personally dedicating my time/resources to resolving.
Not everything is a huge datacenter.
I didn't say they are bad cards... they are just outdated at this point.
If you really want to put your words to action... let me know. I'll put you in touch with someone to buy 130,000 of these cards, and you can sell them to every college kid out there... until then, I wouldn't hold AMD over the coals for not wanting to put effort into something like that when they are already lagging behind on their AI efforts as it is. I'd personally rather see them catch up a bit first.
> I'll put you in touch with someone to buy 130,000 of these cards, and you can sell them to every college kid out there...
Just sell them on eBay? They still go for $50-$100 each, so you're sitting on several million dollars worth of GPUs.
> I'd personally rather see them catch up a bit first.
Growing the community is how you catch up. That doesn't happen if people can't afford the only GPUs you support.
Agreed 100%.
> That doesn't happen if people can't afford the only GPUs you support.
On this part, we are going to have to agree to disagree. I feel like being able to at least affordably rent time on the high end GPUs is another alternative to buying them. As I mentioned above, that is something I'm actively working on.
There are two problems with this.
The first is high demand. GPU time on a lot of cloud providers is sold out.
The second is that this costs money at all, vs. using the GPU you already have. "Need for credit card" is a barrier to hobbyists and you want hobbyists, because they become contributors or get introduced to the technology and then go on to buy one of your more expensive GPUs.
You want the barrier to adoption to be level with the ground.
Something I'm trying to help with. =) Of course, I'm sure I'll be sold out too, or at least I hope so, cause that means buying more GPUs! But at least I'm actively putting my own time/energy toward this goal.
> The second is that this costs money at all, vs. using the GPU you already have.
As much as I'd love to believe in some utopia that there is a world where every single GPU can be used for science, I don't think we are ever going to get there. AMD, while large, isn't an infinite resource company. We're talking about a speciality level of engineering too.
> You want the barrier to adoption to be level with the ground.
100% agreed, it is a good goal, but that's a much larger problem than just AMD supporting their 6-7 year old cards.
I get your point about AMD not wanting to spend money on supporting old hardware, but how do they expect to build a market without a fan base?
Look, I get it. You're right. They do need to work on building their market and they really screwed the pooch on the AI boat. The developer flywheel is hugely important and they missed out on that. That said, we can't expect them to go back in time, but we can keep moving forward. Having enough people making noise about wanting to play with their hardware is certainly a step in the right direction.
https://www.phoronix.com/news/Radeon-RX-7900-XT-ROCm-PyTorch