Open-source project ZLUDA lets CUDA apps run on AMD GPUs
cgchannel.com
cgchannel.com
Noteworthy top comment in that thread:
> This event of release is however a result of AMD stopped funding it per "After two years of development and some deliberation, AMD decided that there is no business case for running CUDA applications on AMD GPUs. One of the terms of my contract with AMD was that if AMD did not find it fit for further development, I could release it. Which brings us to today." from https://github.com/vosen/ZLUDA?tab=readme-ov-file#faq
AMD funded a drop-in CUDA implementation built on ROCm: It's now open-source - https://news.ycombinator.com/item?id=39344815 - Feb 2024 (410 comments)
Zluda: Run CUDA code on Intel GPUs, unmodified - https://news.ycombinator.com/item?id=36341211 - June 2023 (90 comments)
Zluda: CUDA on Intel GPUs - https://news.ycombinator.com/item?id=26262038 - Feb 2021 (77 comments)
Also recent and related:
Nvidia bans using translation layers for CUDA software to run on other chips - https://news.ycombinator.com/item?id=39592689 - March 2024 (155 comments)
I need to put this in all contracts for all jobs in the future.
That said, what _is_ AMD's platform? OpenCL? Vulkan compute? If they don't have an alternative, then the strategy doesn't make sense.
That is a very good question that apparently AMD has no good answer for. I have lost track of the amount of half baked implementations for GPGPU that AMD has attempted and then left to rot. Even if they told me they had the answer now, and they are going to put all their focus on that, I would not trust them to deliver on it. Their best bet is to create implementations for popular libraries like torch that actually stand a chance to work as a drop in replacement.
They heard that it was successful for Google?
A possible solution (that doesn’t involve being better than Nvidia at the things they are good at, which seems to be impossible) is to create frameworks that spit out CUDA or, whatever, OpenCL. Nobody actually wants to use the languages, right? Everybody loves CuBLAS and CuDNN, not CUDA, make GPUOpenBLAS. Maybe they could get Intel to come along with them.
Nobody cares about "users" in this case, it's bespoke applications running on bespoke infrastructure at scale.
"Production grade" feels pretty ambiguous. It either works or it doesn't for any particular developer's use case.
It makes sense in that context.
In EU you could have done it but because of US risks they killed it anyway.
That's not legal but who's gonna stop them.
Also, Wine does the same since forever for DOS binaries. Or NetBSD with compat_* libreries for tons of Unixlike OSes.
Just like it is not legal to do copyright infringement indifferent to how much you own the hardware you do it on.
AFAIK cublas and other first-party libraries are hand-optimized by nVidia for different generations of their hardware, with dynamic dispatch in runtime for optimal performance. Pretty sure none of these versions would run optimally on AMD GPUs because ideally AMD GPUs run 64 threads / wavefront, nVidia GPUs run 32 threads / wavefront.
It is legal for me to make a copy of any copyright protected media using hardware that I own. It is not legal for me to share this copy with others.
Sony Computer Entertainment v. Connectix Corp.
https://scholar.google.com/scholar_case?case=716676913673727...
> The object code of a program may be copyrighted as expression, 17 U.S.C. § 102(a), but it also contains ideas and performs functions that are not entitled to copyright protection. See 17 U.S.C. § 102(b).
> Object code cannot, however, be read by humans.
> The unprotected ideas and functions of the code therefore are frequently undiscoverable in the absence of investigation and translation that may require copying the copyrighted material.
> We conclude that, under the facts of this case and our precedent, Connectix's intermediate copying and use of Sony's copyrighted BIOS was a fair use for the purpose of gaining access to the unprotected elements of Sony's software.
> HIP is very thin and has little or no performance impact over coding directly in CUDA mode.
> The HIPIFY tools automatically convert source from CUDA to HIP.
Nvidia bans using translation layers for CUDA software to run on other chips [1]
____
In this case translating CUDA can allow AMD to chip away at NVidia's market share.
The most famous case of clean-room reverse engineering is for the original BIOS chips back in the early 1980's, where the chips themselves couldn't change- they were hardware! It's going to be orders of magnitude more difficult to do that for software packages that change regularly.
Probably, but Nvidia's market cap suggests there's more than $2 trillion in reasons to front that expense.
While this kind of restriction still isn't legally cut and dry, it does come with tons of precedent, from IBM banning the use of their mainframe OSes in emulation to Apple only licensing OSX for use on Apple hardware.
It's notable however, that both of the examples you give are for Operating Systems, rather than a library which is part of a larger work. Do you know of an example of a single application or library being hardware locked? the only instance I can think of off the top of my head are the old Dos versions of Autocad which had a physical dongle. But even that was DRM and not just an EULA restriction.
Actually, that might be an interesting direction for them to go. Include some key in Nvidia hardware which the library validates. Then they'd get DMCA protection for their DRM.
Pretty much all audio and video production and editing software for many years, and even today. C compilers, for many years, as well. PhysX for awhile. Native Instruments stuff. Saleae logic analyzer software.
Another message in this thread reminded me, too, of the Google Play frameworks on Android, which are also a very good analogy - Google ship these libraries licensed for use only on approved phones.
Those didn't go away in the industry, though AutoCAD moved away from them. Resolume (professional VJ software) and Avid (professional video editing software) still have hardware dongles. Arguably so does Davinci, theirs are just much bigger ;). progeCAD (AutoCAD compatible CAD program) also has USB protection dongles available as a license option.
"You may not reverse engineer, decompile or disassemble any portion of the output generated using SDK elements fo the purpose of translating such output artifacts to target a non-NVIDIA platform."
I don’t think they would have tried that in the 1970s, when there was an antitrust suit against them for disallowing running their software on plug compatible (https://en.wikipedia.org/wiki/Plug_compatible) mainframes.
That (I think) made IBM offer reasonable licensing terms for their software (https://en.wikipedia.org/wiki/Amdahl_Corporation#Company_ori...: “Amdahl owed some of its success to antitrust settlements between IBM and the U.S. Department of Justice, which ensured that Amdahl's customers could license IBM's mainframe software under reasonable terms.”)
That case eventually got dropped in 1982, so it didn’t lead to any jurisprudence as to if/when such restrictions are permitted.
(Aside: for a case that ran for over a decade and produced over 30 million pages (https://www.historyofinformation.com/detail.php?id=923), I find it strange this case doesn’t seem to have made it to Wikipedia yet, and how little there’s elsewhere. Nice example of how bad the public digital record is)
You can't shut down the tools themselves, but you can shut down their use.
https://en.wikipedia.org/wiki/Psystar_Corporation
> On November 13, 2009, the court granted Apple's motion for summary judgement and found Apple's copyrights were violated as well as the Digital Millennium Copyright Act (DMCA) when Psystar installed Apple's operating system on non-Apple computers.
Besides the copyright violation, it is very important to note that the court also considered that circumventing the hardware checks were a violation of the DMCA and illegal in and of itself.
Apple doesn't do anything about the hackintosh ""community"" because they simply don't care about a bunch of random nerds in their basement running macOS but the moment a corporation starts using it to replace their macs you can bet they're going to be sued to oblivion. Not that it would ever happen, hackintosh are going to prove a complete dead end once Apple drops support for x86.
We live in a post-DMCA world. This isn't the era that allowed Bleem to win against Sony, and this is the era that saw the switch emulator developers shit their pants and promise millions to Nintendo in a settlement because they were very unconfident in the possibility of winning in a trial. NVIDIA, for better or worse, has a strong legal standing to clamp down on people who think it would be funny to run their libraries on non-NVIDIA hardware. Do it in your basement if you will, but don't try to push this in a data center.
It doesn't change your point (which is that it appears de facto legally established that IBM can do this), but IBM doesn't completely ban the use of their mainframe OSes in emulation. They are totally okay with people running them in their own emulators (zPDT and ZDT); the thing they won't authorise is people running them on the open source Hercules emulator, since their own emulators cost $$$$, and Hercules is free, and it appears they view the $0 of Hercules as a threat to the $$$$ of their mainframe ecosystem.
In the past they've even authorised third party commercial emulators, such as FLEX-ES. At some point they stopped allowing FLEX-ES for new customers, although I believe some customers who bought licenses when it was allowed are still licensed to use it. But, it isn't impossible they might authorise a third party commercial emulator again – make it expensive, non-open source, and make it only run on IBM hardware (such as POWER Systems), and there's a chance IBM might go along with it.
Emulation is legally protected both explicitly and through legal precedence. The replication of APIs for compatibility purposes has been argued to the US Supreme Court and found to not be copyrightable. At least within some pretty broad scope.
IANAL, but I fail to see what legal basis Nvidia is relying on. For a single user or company who owns no Nvidia hardware this feels moot. For a company with existing Nvidia hardware I could see them having an argument, kinda. But wouldn't that be squarely in the anti-competitive behavior wheelhouse?
Precisely why they are making that statement. The goal is to threaten people who attempt to avoid CUDA
If the CUDA software you want to run on ZLUDA contains first-party Nvidia libraries, which it usually does, you have to care about how those dependencies are licensed.
So in your entire life, you have never downloaded an Nvidia driver and clicked through the EULA? Once you agree, you've agreed.
Or am I totally wrong here?
More EU budget money coming right up.
My joke would be more accurate referring to the EU :-)
Oh, boy.
AMD not developing a competitor to CUDA was the most short-sighted thing I have ever seen. I have no idea why their board hasn't been sacked and replaced with people who understand that you can make the best hardware out there, but if your SW to use it is -- to be very mild -- atrocious, nobody is gonna buy it or use it.
Us, customers, are left to buy the overpriced NVidia cards because AMD's board is too rich to give a damn about a trillion or so of value left on the table. Just... weird. Whoever owns AMD stock I hope is asking questions, because that board needs to go down the nearest drain.
[1] https://github.com/msoos/amdmiscompile -- they eventually fixed this
Indeed. On the flip side they are quite more friendly to open-source, in general, compared to NVidia that's actively hostile(and has been for a while(see Linus "F* you!" video).
Companies that develop hardware generally suck at software. There are exceptions, but they aren't numerous (and indeed have been rewarded in their stock price). I do not know anything about AMD's company culture in their software business units, but fixing that generally requires pretty large changes.
> I have no idea why their board hasn't been sacked and replaced with people who understand that you can make the best hardware out there, but if your SW to use it is -- to be very mild -- atrocious, nobody is gonna buy it or use it.
You probably can't just replace the board (unless C-level mandates are the only thing dragging down the company). You need to replace many more management levels, including a sizable portion of middle management. Sometimes even ICs if software hiring hasn't been properly handled.
True, and you can even get the highest market cap as a hardware manufacturer who suck at software.
those are "community-friendly" segfaults I guess, and it's really only a demonstration of how user-hostile NVIDIA is, what with their working HDMI 2.1 support and compilers and runtimes that actually build and run properly... /s
the "open-source so good!" stuff only really only matters when the open-source stack at least gets you across the starting line. When you are having to debug AMD's openCL runtime or compiler and submit patches because it segfaults on the demo code, that is not "engaging the community as a force-multiplier", it's shipping a defective product and fobbing it off on the community to do your job for you.
It's incumbent on AMD to at least get the platform to the starting line, and their failure to do that has been a problem for over a decade at this point.
Also, honestly, even if you submit a patch how long until they break something else? If demo projects don’t even run… they aren’t exactly doing their technical diligence. People seem to love love love the idea of being an ongoing, unpaid employee working on AMD’s tech stack for them, and nvidia is just sitting there with a working product that people don’t like for ideological reasons…
To wit: the segfault/deadlock issues geohot ran into aren’t just a one-off, they’re a whole class of bug that AMD has been fighting for years in ROCm, and they keep coming back. Are you willing to keep fixing that bug every couple months for the rest of your projects life? After they drop support for your hardware in 6 months, are you willing to debug the next AMD platform for them too?
My naive understanding is that a graphics card is just a funny computer on which you can upload opcodes and data and let it cook itself.
Why is CUDA such a big deal? Can't AMD just give direct access to its GPU as if it was an array of 4096 Arduino boards?
AMD only has CoffeScript and it sucks compared to the TypeScript from NVIDIA.
CUDA is much higher level. It's roughly on par with a "C++ is C with classes" level in terms of language capability. This makes it much easier to develop complex applications. The C compatibility means that you can reuse the exact same code between CPU and GPU in many cases. It eliminates a lot of boilerplate, since you don't need to manage your data in as much detail (eg, while you still have to make sure your pointers are valid for the GPU, the code for uploading the function arguments is generated by the CUDA compiler).
The value add that makes CUDA especially strong is all the first party libraries which have been carefully optimized and have widespread and proven long term support.
If it wasn't taking the tech community so long I would imagine it would not be harder than porting GCC to a new architecture.
The other challenge is that there isn't an "AMD-opcode", each generation tends to change around the opcodes a bit, so you want to compile to an intermediate representation which the driver would ingest and compile down to what the GPU uses. NVIDIA uses PTX for this, it works very well.
AMD's ROCm doesn't use an IR, it compiles the code for each architecture, which means they have a very limited support window for consumer GPUs (to limit binary size). OpenCL and Vulkan supports SPIR-V as an IR, but IIRC, the OpenCL SPIR-V on AMD is very buggy, and Vulkan SPIR-V is very different.
This has other side effects, like BLAS library support range. Hard for an open source community to justify putting in tons of effort into optimizing an entire BLAS library for each generation, when it'll only be supported for ~4 years (so, by the time you're finishing up with driver stability and library optimization, you're probably already at least a quarter of the way through the support period).
Still, giant missed opportunity for AMD not to focus on this when NVIDIA has an almost monopoly on massive GPGPUs.
Amd doesn't. Go watch https://www.youtube.com/watch?v=NPinFkavsrk or https://www.youtube.com/watch?v=AqPIOtUkxNo
That is an attempt to avoid buggy amd implementation and go closer to hw. Everything crashes (kernel), their own demos lock up the card and so on.
It has been a while since I watched these streams, but that is about the theme of all six hours (and few others streams). It's just complete mine field, where he is trying to walk through a very narrow path of success.
Combine it with things like
this (they now actually have a improved documentation, significant progress): https://github.com/ROCm/ROCm/issues/1714
or this (HIP doesn't support L2 cache coherency, so disable GPU L2 cache) https://docs.amd.com/projects/HIP/en/latest/user_guide/hip_p...
and it's just FUBAR.
I actually rather enjoyed the AMD Software in particular, since it made very easy to tweak graphics (limit framerates to 60 when I don't want the GPU maxing out when games/software don't support it by default), setup instant replays with a hotkey press (like Shadowplay, where it has a constant recording buffer of the last X minutes) and also both power limit the GPU (when my UPS wasn't very good) as well as overclock it automatically (since I still want to squeeze like a year out of my RX 580).
Except that any version of the software/drivers after around 2020 crashes VR titles after less than an hour. And that there is no software package for Linux and CoreCtrl isn't as good. And that sometimes the instant replay thing just doesn't work. And that I haven't been able to get ROCm working with any of the local LLMs even once across both Windows and Linux (DKMS sure loved to do a whole bunch of pointless compiling upon each apt upgrade).
I'm honestly considering either going for Intel Arc as my next GPU because I'm curious, or just going back to something from Nvidia, so it's probably a split between: A580, RX 6600, RTX 3050. Or maybe I can hold out until other parts drop in price, time will tell.
Is AMD trying to create their own proprietary lock in eco system? Why aren't they committing to cross platform open standards?
Without CUDA, there's absolutely no way NVIDA would be a nearly trillion dollar company today.
I tried ROCm on my iGPU last year and you do get a bit of a benefit for prompt processing (5x faster) but inference is basically bottlenecked by the memory bandwidth whether you're on CPU or GPU. Here were my results: https://docs.google.com/spreadsheets/d/1kT4or6b0Fedd-W_jMwYp...
Note, only GART memory, not GTT is accessible in most inferencing options, so you will only basically be able to load 7B quantized models into "VRAM" (BIOS controlled, but usually max out at 6-8GB). I have some notes on this here: https://llm-tracker.info/howto/AMD-GPUs#amd-apu
If you plan on running a 7B-13B model locally, getting something like a RTX 3060 16GB (or if you're on Linux, the 7600 XT 16GB might be an option w/ HSA_OVERRIDE) is probably your cheapest realistic option and will give you about 300 GB/s of memory bandwidth and enough memory to run quantizes of 13/14B class models. If you're buying a card specifically for GenAI and not going to dedicate time to fight driver/software issues, I'd recommend going w/ Nvidia options first (they typically actually give you more Tensor TFLOPS/$ as well).
Such a dick move from AMD.
Idk, running 7B language models on my 6650 XT with ROCm has been pretty slick. Doesn't seem like fraud to me.
Compile each one to get a binary.
Train a language model with the source and output binary.
Hey presto, clean room compiler.
Edit: Oh wait.. duh.. just train it on the equivalent target source.
Presumably you can do this for other targets as well.