Zluda: Run CUDA code on Intel GPUs, unmodified
github.com
github.com
Longer term, I think we can write GPU code that is portable, but it will require building out the infrastructure for it. Vulkan compute shaders are one good starting point, and as of Vulkan 1.3 the "subgroup size control" feature is mandatory. WebGPU is another possible path to get there, but it's currently lacking a lot of important features, including subgroups at all. There's more discussion of subgroups as a potential WebGPU feature in [1], including how to handle subgroup size.
Then they go to the developers and ask why the implementation isn't optimized for this hardware lots of people have and the solution is to do an implementation in Vulkan etc.
NVIDIA used to discourage code which relies on the subgroup or warp size. I'm not sure how much this is true of real world code though.
Both Intel and AMD have the opportunity to create some actual competition for NVIDIA in the GPGPU space. Intel, at least, I can forgive since they only just entered the market. Why AMD has struggled so hard to get anything going for so long, I don't know...
I shouldn't be surprised that AMD's ecosystem is lagging behind, since their GPU division spent a good decade suffering to be even relevant. Not to mention that NVIDIA has spent a lot of effort on their HPC tools.
I don't want this to be too negative towards AMD, they have been steadily growing in this space. Some things do work well, e.g. stable diffusion is totally fine on AMD GPUs. So they seem to be catching up. I just feel a little impatient, especially since their cards are more than powerful enough to be useful. I suppose my point is that the gap in HPC software between NVIDIA and AMD is much larger than the actual capability gap in their hardware, and that's a shame.
While I am a little bit of a fan of AMD, there's still work to do. I think AMD really needs to take advantage of their production margins to gain more market share. They also need to get something a bit closer to the 4090 on a performance gpu + entry workstation api/gpgpu workload card. The 7900 XTX is really close, but if they had something with say 32-48gb vram in the sub-2000 space it would really get a lot of the hobbiest and soho types to consider them.
Their hardware is good but if all they were selling was Macbooks and iPhones with Windows and Android on them, they wouldn't have anything near their current margins.
If they hadn’t made platform changes they would have never been able to turn into what they are today. I hardly thing that is ‘little to do’.
They would likely barely exist. They have ‘achieved product market fit’ as the saying goes. Which requires more than just a sharp UI, as their history shows
It was more forced upon them than anything else.
The move to M1 was an actual innovation that came after their success.
The real story of the company though is the iPhone, which is absolutely their own technical innovation.
They are a brand first, and a technical company second. That doesn't mean they aren't doing cool technical things. But a lot of companies do cool technical things and still fail.
I really hope AMD manages a comeback like they showed a few years ago with their CPUs. Intel joining the market is certainly helping, but having three big players competing would certainly be desirable for all sorts of applications that require GPUs. AMD cars like the 7900 XTX are already fairly promising on paper with fairly big VRAMs, they'd probably be much more cost effective than NVIDIA cards if software support was anywhere near comparable.
[0]: https://www.amd.com/en/graphics/servers-solutions-rocm
[1]: https://pytorch.org/
While cuda works on a 1050ti.
To get an idea, the 1050ti is a card with an MSRP of $140 - almost seven years ago when it was released. Between the driver and CUDA support matrix it will likely end up with a 10 year support life.
While it's not going to impress with an LLM or such it's the lowest minimum supported card for speech to text with Willow Inference Server (I'm the creator) and it still puts up impressive numbers.
Same for the GTX 1060/1070, which you can get with up to 8GB VRAM today for ~$100 used. Again, not impressive for the hotter LLMs, etc but it will do ASR, TTS, video encoding/decoding, Frigate, Plex transcoding, and any number of other things with remarkable performance (considering cost). People also run LLMs on them and from a price/performance/power standpoint it's not even close compared to CPU.
The 15 year investment and commitment to universal support for any Nvidia GPU across platforms (with very long support lifecycles) is extremely hard to compete with (as we see again and again with AMD attempts in the space).
When comparing the last driver release from Nvidia in the 2014-205 range it supports cards going back to at least 2004[1].
[0] - https://download.nvidia.com/XFree86/Linux-x86_64/390.157/REA...
[1] - https://download.nvidia.com/XFree86/Linux-x86_64/96.43.23/RE...
And failed to make any of them work, which to my mind means they've burned their possibilities more than if they flat-out did nothing.
Apple Silicon being on Metal Performance Shaders (I think they deprecated OpenCL support?) kind of makes this all more confusing.
It definitely feels like CUDA is the leader and anything else is backseat/a non starer, which is fine. The community support isn't there.
I haven't heard anybody talk about AMD Radeon GPUs in a looong time.
So it is already a non starter if they can't meet those baselines.
A note to any systems hackers who might try something like this, you can also retarget clang's CUDA support to SPIR-V via a LLVM-to-SPIR-V translator. I can say with confidence that this works. :)
So no then.
"What is the status of the project?
This project is a Proof of Concept. About the only thing that works currently is Geekbench. It's amazingly buggy and incomplete. You should not rely on it for anything serious."
it works 100% of the time. until it does not.
"It's compatible with CUDA as long as you don't use all the features of CUDA."
So it's a drop in replacement for some subset of modern CUDA. I feel like most folks who are upvoting this don't program CUDA professionally or aren't very advanced in their usage of it.
“Complete” = it covers everything.
It is drop-in, but not complete.
I think that what you're trying to say is that before claiming to be a "drop-in replacement", make sure that your supported feature set is representative enough of mainline CUDA development.
I think this might get better if and when people redesign their dev workflows, CI/CD pipelines, builds, etc, to deploy code to both hardware platforms to ensure matching functionality and stability. I'm not going to hold my breath just yet. But it would be really great to have two viable platforms/players in this space where code can be run and behave equally.
I had to write a cross-platform kernel a few weeks ago, and I ended using pre-processor guards to make it work with the OpenCL and CUDA compilers [1].
[1] https://github.com/RaphaelJ/libhum/blob/main/libhum/match.ke...
Zluda: CUDA on Intel GPUs - https://news.ycombinator.com/item?id=26262038 - Feb 2021 (77 comments)
However, they have gone around that by creating HIP which is a CUDA adjacent language that runs on AMD and also translates to CUDA for Nvidia GPUs. There is also the HIPify tool to automatically convert existing sources from CUDA to HIP. https://docs.amd.com/bundle/HIP-Programming-Guide-v5.3/page/...
Furthermore, CUDA is a language (dialect of C/C++) not an API, so that precedent may not have much weight.
A more aggressive approach was tried during the Xbox 360 era with the Games For Windows Live framework and by removing their games from the Steam store. It ended up catastrophically bad and they had to backtrack on both decisions.
The irony of Proton de facto killing any chance for native linux ports of windows games isn't lost to them, either.
Playing games on linux is not a threat to microsoft. The money they loose on that is miniscule.
I'm reasonably sure they could reimplement CUDA from a copyright / trademark perspective. It's possible that they could be blocked with patents though.
IIRC, the verbatim copying of rangeCheck didn't make it to SCOTUS. They really did instead rule on the copyrightability of the "structure, sequence, and organization" of the Java API as a whole.
You can guess how much money lawyers have been paid over that circumstance.
And FWIW, that seems to be a reasonable result given the overall market structure at the moment. Having all eggs in the Nvidia basket is great for Nvidia shareholders, but not for customers and probably not even for the health of the surrounding industry.
Let's suppose that an open-source CUDA API is in a legal gray zone that could only be clarified by a judge.
Could a company like AMD create a wholly owned subsisiary to make an attempt, whithout exposing the parent company to legal liability?
> Is ZLUDA a drop-in replacement for CUDA?
> Yes, but certain applications use CUDA in ways which make it incompatible with ZLUDA
> What is the status of the project?
> This project is a Proof of Concept. About the only thing that works currently is Geekbench. It's amazingly buggy and incomplete. You should not rely on it for anything serious
It is a cool proof of concept but we don’t know how far away it is from becoming something that a company would willingly endorse. And I suspect AMD or Intel wouldn’t want to put a ton of effort into… helping people continue to write code in their competitor’s ecosystem.
https://www.jwz.org/blog/2012/06/i-have-ported-xscreensaver-...
https://dereferer.me/?https%3A//www.jwz.org/blog/2012/06/i-h...
Went there a few days ago. Got a colonoscopy picture. Even without deferrer/referrer.
I failed to find information regarding Fortran, Haskell, Julia, .NET, Java support for CUDA workloads.
Running CUDA on an AMD RDNA3 APU is what I'd like to see as its probably the cheapest 16GB shared VRAM solution (UMA Frame Buffer BIOS setting) and creates the possibility of running 13b LLM locally on an underutilized iGPU.
Aaand its been dead for years, shame.
- AMD already has a CUDA translator: ROCM. It should work with llama.cpp CUDA, but in practice... shrug
- Copies the CUDA/OpenCL code make (that are unavoidable for discrete GPUs) are problematic for IGPs. Right now acceleration regresses performance on IGPs.
Llama.cpp would need tailor made IGP acceleration. And I'm not even sure what API has the most appropriate zero copy mechanism. Vulkan? OneAPI? Something inside ROCM?
But… I don’t know if it’s possible to do for iGPUs that partition memory in BIOS. I am curious for the answer.