VUDA: A Vulkan Implementation of CUDA
github.com
github.com
> It all just works. No code changes were needed.
To make it even more insulting, even simply installing ROCm itself is a massive burden, even on an ostensibly-supported (as geohot discovered) and even just “it works out of the box if you distribute and compile it locally” is ignoring that whole massive “draw the rest of the owl” stage of getting ROCm installed and building properly in your environment.
Even a source-compatible layer that let you just recompile CUDA code for an AMD GPU would be a huge improvement. That alone would eliminate the CUDA lock-in.
I have enough issues using graphics already so I'll stick with Mesa.
But it currently only runs on CDNA boards and enterprise-y Linux distros (Ubuntu LTS, Centos, etc.)
TLDR; If you provide even more functions through the overloaded headers, incl. "hidden ones", e.g., `__cudaPushCallConfiguration`, you can use LLVM/Clang as a CUDA compiler and target AMD GPUs, the host, and soon GPUs of two other manufacturers.
FWIW, we have some alternative ideas on how to get out of the vendor trap, as well as some existing prototypes to deal with things like CUBLAS and Thrust. Feel free to reach out, or just keep an eye out.
2. Implementing the _runtime_ API is not the right choice; it's important to implement the _driver_ API, otherwise you can't isolate contexts, dynamically add newly-compiled JIT kernels via modules etc.
3. This is less than 3000 lines of code. Wrapping all of the core CUDA APIs (driver, runtime, NVTX, JIT compilation of CUDA-C++ and of PTX) took me > 14,000 LoC.
What I _really_ like to receive, though, is feedback from using the wrappers, ideas for changes/improvements, and of course messages volunteering to QA new versions before their release :-P
I'm not an expert here but this approach seems powerful and important. But this system seems complex enough to doubt the ability of an individual to build. It seems like this would need a corporate sponsor to get off the ground. Perhaps AMD itself would be interested in paying engineers to iterate on this?
> The software is terrible! There’s kernel panics in the driver. You have to run a newer kernel than the Ubuntu default to make it remotely stable. I’m still not sure if the driver supports putting two cards in one machine, or if there’s some poorly written global state. When I put the second card in and run an OpenCL program, half the time it kernel panics and you have to reboot.
He also talks about user space stuff but clearly he thinks the whole stack, above and below this kind of library also needs a lot of work.
I'm sure even if their GPUs were twice as fast as Nvidia's, everybody would still buy team green because it's better to have a card that works than a broken piece of garbage. We tried to get an MI50 to work reliably at work with KVM, but that thing was a complete dumpster fire. A colleague just bought a 7900XTX for gaming and spent days getting it to work. This included three Windows reinstalls. And that use case is gaming on Windows, which supposedly is the best supported case. It only gets worse from there. Compute on Linux? Lol.
Now last time this topic came up, someone claimed that AMD is pretty much at their limits production wise, and there are a few unnamed large companies buying loads of their cards for compute and cloud gaming, and AMD basically has engineers dedicated to making sure things work exactly for their use case, so they don't have to really care about the rest. Sounds pretty wild, but not completely unrealistic...
My old nvidia 570's drivers went into severe bitrot. Basic stuff like screensavers and desktops broke badly, and games were flaky. The card is still more than powerful enough for what I used it for.
I switched to AMD, with open source drivers. I get windows-level performance on AAA (and indie) games in steam, and zero compatibility issues with the rest of the Linux ecosystem.
I'm considering switching back to an AMD card due to this.
At least Nvidia finally open sourced their driver too, I guess. And Intel is still open source. But it still sucks a bit I think unless you do research.
I’m being a little tongue-in-cheek here, but the best supported case for AMD is gaming via console: AMD provides CPU/GPU for the current generation of both the XBox and PlayStation consoles.
Which suggests to me that they shouldn’t have too much problem supporting their hardware on Windows or Linux. But that’s outside of my area of expertise. Maybe they need to spend too much engineer effort and time supporting the consoles at what’s probably a pretty thin profit margin?
Oh, of course.
You can't seriously tell me that's not something they could fix.
https://www.youtube.com/watch?v=Mr0rWJhv9jU
and
https://geohot.github.io/blog/jekyll/update/2023/06/07/a-div...
I feel a lot better about my journey with AMD now; there seemed to be some major issues with their GPU drivers. Now I know it wasn't just me.
cuDNN being higher level offers more opportunity for compatibility without losing performance (i.e different implementations of kernels fine-tuned for optimal performance on AMD vs NVidia hardware), but the trouble is that so much of what frameworks like PyTorch do is based on custom kernels, not just cuDNN.
It seems the best bet for AMD would be a rock solid low level API (not a moving target) and support of high level optimizing ML compilers to reduce the level of effort for the framework (PyTorch, TensorFlow, JAX ...) vendors to provide framework-level support on top of that. Ultimately they'd need to work very closely with the framework vendors to provide this support, since they are the ones who would be benefiting from it.
It's odd how much of an afterthought ML support has seemed to be for AMD over the years... maybe the relative size of the consumer ML market vs graphics/gaming market didn't seem to make it worth their effort, but as NVidia has shown this is a path to gaining much more lucrative data center wins.
>>> hipify-clang is a clang-based tool for translating CUDA sources into HIP sources. It translates CUDA source into an abstract syntax tree, which is traversed by transformation matchers. After applying all the matchers, the output HIP source is produced. [...]
(Edit) CUDA APIs supported by hipify-clang: https://rocm.docs.amd.com/projects/HIPIFY/en/latest/supporte...
https://en.wikipedia.org/wiki/3dfx_Interactive#Product_devel...
I can’t see NVIDIA letting this just exist
The situation is starting to improve though. Installed a bunch of libraries from https://repo.radeon.com/rocm/apt/5.4 jammy main and the crashes got less frequent. I don't have a lot of faith in AMD to deliver reliable BLAS libraries at this point, but it could happen. The hardware is there, I just don't think they're prioritising supporting the right places in the distribution chain or supporting consumer-level graphics.
The only saving grace would be Oracle v Google which established the de jure that an API isn't copyrightable.
The main concern with Oracle v Google was that the court would ignore or misinterpret the existing precedent.
A secondary concern was that a Google employee formerly worked on Java at Sun (and/or Oracle), and copy-pasted some implementation source code from oracle to google's code bases. There was a real possibility the "APIs aren't copyrightable" precedent would stand, but the courts would rule that Google couldn't continue distributing Dalvik.
How far is it actually compatible right now?
Are there any tests / benchmarks?
Can this be used to run CUDA-accelerated LLMs?
So the answer is no, it can't be used with kernels that use cublas or cudnn, which excludes almost all ML use-cases.
https://registry.khronos.org/vulkan/specs/1.3-khr-extensions...
However, the deep learning field does currently not pay much attention to reproducibility, so this might not be a big issue.