NVIDIA could voluntarily open the standard to avoid this ligitation if they wanted to, though, and IMO it would be the smart thing to do, but almost every corporation in history has chosen the ligitation instead.
NVIDIA could voluntarily open the standard to avoid this ligitation if they wanted to, though, and IMO it would be the smart thing to do, but almost every corporation in history has chosen the ligitation instead.
If AMD isn't a competitor before government intervention, I don't the government forcing nvidia to open up CUDA changes much. CUDA's moat isn't due to some secret sauce - nvidia put in the developer hours; and if AMDs CUDA implementation is still broken, people will continue to buy nvidia.
There has been a lot of trying to get AMD to work - Hotz has been trying for a while now[1] and has been uncovering a ton of bugs in AMD drivers. To AMD's credit, they have been fixed, but it does give you a sense of how far behind they are in regards to their own software. Now imagine them trying to implement a competitor's spec?
[1] https://twitter.com/__tinygrad__/status/1765085827946942923
The issue that AMD has is they had a long period where they clearly had no idea what they were doing. You could tell just from looking at websites, CUDA pretty much immediately gets to "here is a library for FFT", "here is a library for sparse matricies". AMD would explain that ROCM is an abbreviation of the ROCm Software platform or something unspeakably stupid. And that your graphics card wasn't supported.
That changed a few months ago; so it looks like they have put some competent PMs in the chair now or something. But it'll take months for the flow on effects to reach the market. They have to figure out what the problems are which takes months to do properly; then fix the software (1-3 months more minimum); then get it into the open and the foundational libraries like PyTorch pick it up (might take another year). You can speed that up, but more cooks in the kitchen is not the way. Bandwidth use needs to be optimised.
It isn't like ROCm seems lacks key features; it can technically do inference and training. My card crashes regularly though (might be a VRAM issue) so it is useless in practice. AMD can check boxes but the software doesn't really work and grappling with that organisationally is hard. Unless you have the right people in the right places, which AMD didn't have up to at least mid 2023.
If even that. A few years ago they managed to break basic machine learning code on the few commonly-used consumer GPUs that were officially supported at the time, and it was only after several months of more or less radio silence on the bug report and several releases that they declared those GPUs were no longer officially supported and they'd be closing the bug report: https://github.com/ROCm/ROCm/issues/1265
It makes perfect sense that, organisationally, they were focused on that battle. If you remember the Athlon days, AMD beat Intel before, but briefly. It didn't last. This time it looks like they beat Intel and have had the focus to stay. Intel will come back and beat them some cycles, but there is no collapse on the horizon.
So it makes sense that they started looking at nVidia in the last year or so. Of course nVidia has amassed an obscene war chest in the meantime...
Pytorch has been "supporting" rocm for all last 2 years
If they paid everyone $1M/year salaries maybe more people would consider going into aerospace engineering.
Right now though Boeing's starting salaries aren't that much higher than what an Uber driver in the bay area makes.
The problem currently, as people like Hotz and many others are discovering, it not the lack of CUDA. Most people use PyTorch and don't care what the underlying software is. Infact most CUDA is hand tuned to nvidia hardware anyways and is optimized to make the most on nvidia. The problem is AMD's drivers - the piece that actually sends the code to run on the GPU, tends to be broken. AMD cannot "sponsor" an outsider to fix this. A legal, but broken, AMDCUDA will not be any better than the current situation; so no, having CUDA on AMD wouldn't change anything.
The problem is not "CUDA is not AMD", the problem is AMD has not, does not, and for some reason will not invest adequately in GPU compute. CUDA is a mirage; if AMD had a similar platform someone would have done the work already to ensure PyTorch works on it. PyTorch already supports ROCm, people don't use it because the performance is bad and it's buggy. When nvidia had this problem, nvidia hired engineers to work on open source projects and debug issues in open source libraries (not even limited to AI, you will find nvidia engineers debugging issues in a wide range of CUDA projects). When AMD has this issue, they barely acknowledge it.
If it was a clean room implementation of the API NVIDIA wouldn’t care. Heck that’s exactly what AMD did with HIP.
But what you cannot do is essentially intercept calls to and reverse engineer NVIDIA binaries in real time because you can’t be arsed to build your own.
And this is precisely what anti-trust ligitation would allow them to do.
Preventing someone from reverse-engineering a product with the sole intention of maintaining monopoly status may be seen as anti-competitive.
AMD makes really, really good CPUs now, but only after ligitation against Intel allowed them to keep up with evolving x86 standards.
It's not about being "arsed" to build your own, the problem is NVIDIA controls the ecosystem-wide standard. NVIDIA can add to CUDA at any point in time and launch a GPU at the same time, the ecosystem would be forced to buy it if they want to stay on the cutting edge, and AMD would never be able to compete or reverse engineer these new standards in time.
What ZULDA did wasn’t to maintain a compatibility with the CUDA API and provided an open implementation of it but rather use all the CUDA based libraries that NVIDIA provided on top of it.
The equivalence again would be that not only Google implemented their own Java compatible API but rather that they used the now Oracle owned JVM to do so and redistributed it.
AMD already implemented CUDA essentially one to one in the form of HIP in ROCm. The issue that they face is that they don't have all the equivalent middleware to make copy pasting code actually work and this was what ZULDA did but instead of building a HIPdnn they just reused NVIDIA binaries.
The MI300 smokes the H100 yet here we are.
CUDA is such a misnomer. Amd doesn't have tensorRT, cuDNN, cutlass, etc. Forcing Nvidia to make these work on AMD is like forcing Microsoft to make windows work on apple hardware... Not going to happen.
You can absolutely use the names of the functions and the programming model. Like I said, HIP is literally a copy. Llama.cpp changes to HIP with a #define, because llama.cpp has its own set of custom kernels.
And this is what I've said before, CUDA is hardly a moat. The API is well-known and already implemented by AMD. It's all the surrounding work: the thousands of custom (really fast!) kernels. The ease-of-use of the SDKs. The 'pre-built libraries for every use case'. You can claim that CUDA should be made open-source for competition, but all those libraries and supporting SDKs represent real work done by real engineers, not just designing a platform, but making the platform work. I don't see why NVIDIA should be compelled to give those away anymore than Microsoft should be compelled to support device driver development on linux.