There's still work to do -- lots of repositories still contain `if device=="cuda"` type language, but my own experience is that manging code around to use apple gpus has gotten vastly easier this year, and I see more and more AMD GPU owners floating around github issues with resolvable problems. A year ago they were barely present.
All that said - people aren't ignoring it - and entrepreneurs and cloud monopolies are putting real resources into opening things up. I think the playing field will continue to level / get back to competitive over the next two to three years.
And while many problems are trivial to fix at eg the PyTorch layer, lots of stuff like flash attention, DeepSpeed, etc. are coded directly in CUDA kernels.
How is CUDA so sticky when most ML devs aren't writing CUDA but something several layers of abstraction above it? Why can't intel, AMD, google w/e come along and write an adapter for that lowest level to TF, pytorch or whatever is the framework of the day?
A long long time ago, i.e. the last time AMD was competing with Intel (before this time, that is), we used to use Intel's icc in our lab to optimize for the Intel CPUs and squeeze as much as possible out of them. Then AMD came out with their "Athlon"(?) and it was an Intel-beater at that time. But AMD never released a compiler for it; I bet they had one internally, but we had to rely on plain old GCC.
These hardware companies don't seem to get that a kick-ass software can really add wings to their hardware sales. If I were a hardware vendor, I would, if nothing else, make my hardware's software open so the community can run with it and create better software; which will result in more hardware sales!
This is a really tricky guess, case in point that AMD's latest chip cant compete on training because they could not get Flash Attention 2 working on the backward pass because of their hardware architecture. [1]
Attempts to abstract at a higher layer have failed so far because that lower layer is really valuable, again Flash Attention is a good example.
[1] https://www.semianalysis.com/p/amd-mi300-performance-faster-...
The API that matters is Torch, and only the API. Letting NVIDIA charge famine prices is both a bad idea and a huge incentive to write software bridging the gap.
The amount of extremely CUDA specific handwritten compute architecture code, custom kernels, etc has exploded and other than a few things here and there (yes, like torch SDPA) we’re waaaay past vanilla torch for anything beyond a toy.
The pricing and theoretical hardware specs of these novelties doesn’t matter when an inferior on paper Nvidia product will wipe the floor with the other in the real world.
People have spent 15 years wringing every last penny out of CUDA on Nvidia hardware.
There is some light shining through but you’ll still see things like “Woo-hoo FlashAttention finally supports ROCm! Oh wait, what’s that, FlashAttention2 has been running on CUDA for six months?”
Don’t even get me started on the “alternative” software stacks and drivers.
The graphical debugging tools that allows stepping through GPU code just like with on the CPU.
This is why everyone is trying to compete with CUDA. Everyone wants a slice of the pie as we go from tiny amount to everything.