https://www.amd.com/en/products/software/rocm.html
Before that, they were pushing heavily for standard OpenCL, but that failed because the hardware wasn't as competitive and the ecosystem/tooling barren.
Yes, “just use libraries”. But as Andrej Karpathy used to say, “I don't need some library holding my hand and providing abstractions. Real men command GPUs with their own raw kernel code”.
(And, incidentally, there are much less libraries for ROCm, for this reason. Somebody should write them, and why do that with something that takes 5x more code to do the same thing?)
OK, I retract my statement.
To get a picture of the current state which has changed a lot this MS Ignite presentation may be of interest. https://youtu.be/7jqZBTduhAQ?t=61
And there's MI300A which has a combined CPU+GPU with shared memory space which already has supercomputer customers. You can say it's DOA relative to Hopper but MI300 will be AMD's most successful GPGPU yet.
If they can’t afford to out R&D Nvidia today, how could they afford to fund a competitor? In the past 20 years the only serious new entrant to the discrete GPU market has been Intel and they are also far behind Nvidia and CUDA.
I was actually motivated to contribute to the AMD drivers at some point and did land a couple of patches, but wasn't looking to switch career paths. The drivers have become quite good without me anyway, so no regrets :)
At the very least, pytorch has to work effortlessly across GPUs, down to the less popular and odd workarounds people build into their pyotrch code.
Even if RocM reaches parity, the first mover advantage for Cuda is too large. RocM porting has to be literally effortless. Everything else is DOA.
That being said, if you a cloud provider and need to scale up a bunch of basically similar transformers models, then it should be an easy sell for AMD.
And the large companies (Meta, Microsoft) buy these GPUs by the thousands upon thousands.
Having a small team of engineers spend a couple months porting code over is well worth it for even modest cost reductions.
These may not be useful to smaller companies working in the AI space, but odds are to sellout, all AMD really needs are the sales contracts that they've already announced.