Not true. PyTorch has had an MPS backend[1] for Apple Silicon for quite some time now. And lately, they've also added a HIP backend[2] for AMD hardware. I've written CUDA kernels for PyTorch in the past, it'd have been a lot easier if PyTorch were a thin wrapper over CUDA, but it isn't.
==> You can also say it's a lot of other things like ROCm. <==
I mean my intent was just to communicate that PyTorch runs CUDA in the backend, which was what I meant by API, because it seemed the person I was responding to wasn't aware of this backend code you're interfacing with. I thought adding the same phrasing about it also "being" ROCm and using "such as" and "other things," that others would have allowed anyone with more knowledge to infer that I'm aware of the generalization but I guess I was deeply mistaken given the replies that are certain that I can't read the pytorch homepage. Though I'm not quite sure how you came to the conclusion that I wasn't aware of ROCm support since I did explicitly mention it.
Do you have feedback for how I can better communicate? I seem to be running into failures like this a lot lately where I feel like I can point to where I clearly stated something but that doesn't matter because communication is about getting the person on the other side to receive the message. I've been scratching my head about this and could definitely use the outside perspective.
Tensorflow had first mover advantage but they also had first mover disadvantage. TF isn't that easy to write nor to debug. PyTorch got to see everything that was wrong with TF and fix it. No doubt Meta has put in lots of amazing effort. It is an incredible piece of software.
I'll also add that PyTorch does a lot of great stuff that people don't use or recognize. There are a lot of good distributional works out there for prob and stats people that give you cuda acceleration. Really a lot of numpy can be replaced with pytorch and its great. But be warned of FMA and other optimization differences.
Obviously you can use cuda to write more than ML algos, but that’s its overwhelming use case.
Thus my point is if AMD ported PyTorch to their hardware (and PT is not cuda-only) it would create a more level playing field on which they could compete (in a cousin vs cousin battle) by eliminating the switching cost for the majority of developers.
This is the same reason cpu mfrs have compiler ports but generally don’t have to worry about instruction set level compatibility.
It is ported (I think by AMD). I use it daily. Biggest problem is that the people maintaining PyTorch aren't very interested as they don't use it themselves. Hence there are silly bugs in it. Users contribute fixes and they get closed down as "unsupported" or get plain misunderstood.