Zluda: CUDA on Intel GPUs
github.com
github.com
Growing up, people sometimes called their kids "zludy" (plural of "zluda" in my language) -- or trouble-makers.
But a look at HW Acceleration support table at FFmpeg[2] shows why GPGPU Platform API is such a mess. But performance benefits are incredible, using VAAPI for FFmpeg to encode 1080p 2560x1080 screen capture at 60fps reduces CPU usage from 90% to 10% on a old corei5 with intel HD 3000; An old laptop could be perfectly used as an encoding machine for streaming just by using HW Acceleration.
What's funny is that the laptop also has Radeon HD 6490M with 1GB GDDR5 dedicated memory and it's not supported by VAAPI for encoding! GPGPU API/Platform Support are astonishingly messy.
Doesn’t this prove the broader point that GPGPU cross platform would benefit everyone? A new codec is written.. and everyone gets to use it not just those with fixed function support.
CUDA(NVENC/NVDEC), AMF, OpenCL are listed on the FFMpeg chart I mentioned and linked to. HW Accel might use different compute unit for its function, but GPGPU API is also being used for that nevertheless.
Perhaps HW Accel is not the best example for sighting the need gap for universal GPGPU API or its implementation as OP talks about porting GPGPU API for a non-supported platform; What would be the right example?
P.S. I stand corrected on implying Radeon HD 6000 couldn't encode in HW Accel due to API mess, where as it was missing ASIC. Thanks.
The only GPGPU library in that chart is OpenCL, used for accelerating custom video filter effects like blur. The table says OpenCL is supported on Linux, Windows, Mac, and Android devices (with GPUs capable of GPGPU) for Intel, AMD, and Nvidia. The Pi isn't supported because the GPU isn't capable enough. That leaves the only capable platform not covered by that could be as iOS (walled garden reasons). I.e. I'm not seeing the gap you keep referring to for GPGPU APIs, just a task (video coding) that doesn't work well on general purpose hardware.
As for ZLUDA, why is it needed then? CUDA was and is extremely popular as Nvidia has been at the forefront of GPGPU (hardware and software) for over a decade pushing CUDA along the way. OpenCL didn't even get comparable until 5 years after CUDA kicked off. As a result there is a large amount of CUDA code out there that people don't want to rewrite just to make it portable. Hence projects like ZLUDA for Intel which does it transparently and ROCm from AMD which helps automate the porting of code.
Not necessarily, NVENC can use CUDA cores for hardware acceleration too for certain features[1] -
>HEVC encoding also lacks Sample Adaptive Offset (SAO). Adaptive quantization, look-ahead rate control, adaptive B-frames (H.264 only) and adaptive GOP features were added with the release of Nvidia Video Codec SDK 7.These features rely on CUDA cores for hardware acceleration.
>Nvidia Video Codec SDK 8 added Pascal exclusive Weighted Prediction feature (CUDA based). Weighted prediction is not supported if the encode session is configured with B frames (H.264).
Without looking into the CUDA licenses (who knows, they might even expressly allow this kind of thing, but seems pretty unlikely to me), I'd expect this to be a case of whether "APIs are copyrightable" or not, same as the famous and sort of still ongoing https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
The US courts said "yes" in this case (note: 100% stupid IMO), but I'm not confident that nVidia'd have an easy win if they decided to sue the developer, and I'm also relatively sure they wouldn't send a DMCA request at this stage (and that if they did, their request would probably be reviewed harder than normal).
It is not usual for CAFC to hear copyright disputes; that it was appealed to CAFC instead of the 9th Circuit is because there was a patent claim at one point, and CAFC should have followed 9th Circuit precedent on the matter. Google contends that 9th Circuit precedent holds that the API is not copyrightable, which means that CAFC erred in ignoring precedent. Most software companies ultimately agree with Google here, not Oracle: it's telling that most of the amici who side with Oracle are not software companies but media publishers (e.g., MPAA, RIAA).
https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
Edit: huh? I pasted it in with the period and it gets removed. Do trailing periods get removed because it thinks it's the end of a sentence?
Anyway, this redirects to the right URL: https://en.wikipedia.org/wiki/Google_v_Oracle
https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_...
So long as Nvidia's case doesn't get outright laughed out of court, they can throw money at their legal team until the developer goes bankrupt.
At one of the internal Q&As, there was a question as to why we didn't just implement CUDA. One of the reasons given was that the lawyers looked at the license of CUDA and decided that Intel could not legally implement CUDA for Intel's GPU devices. I don't know the details, but quite frankly, it wouldn't surprise me if Nvidia didn't somehow put a poison pill in there to prevent Intel or AMD from implementing it for their own GPUs (note that AMD also doesn't provide an implementation of CUDA for its own GPUs).
Instead, the strategy Intel pursued was to develop a migration tool from CUDA to Sycl: https://software.intel.com/content/www/us/en/develop/tools/o...
Here is how you beat the CUDA lock-in: consistently make better performing GPUs so not using them is a liability. Instead buying AMD you not only get a worse GPU but also the intern software solution, and that is just not compelling.
Also they keep forgetting that CUDA is not only C++, rather a polyglot eco-system.
Instead of targeting an IR, it directly targets a given GPU's ISA, so that your existing binary will not run on future hardware. That's a total no-go for basically every non-HPC use case.
Intel is much better off building a sound technical foundation from scratch.
With oneAPI I had hoped to get the inverse, a oneAPI implementation for NVIDIA hardware, but I don't think the CUDA driver API is low-level enough to do so (e.g. explicit vs global contexts). And yes, I know of Codeplay's implementation of DPC++ for NVIDIA GPUs, but that doesn't implement oneAPI Level0 APIs so is not usable for other languages.
I fail to see how that is a drop-down replacement.
Plus apparently it doesn't support the polyglot CUDA ecosystem.
> Warning: this is a very incomplete proof of concept. It's probably not going to work with your application. ZLUDA currently works only with applications which use CUDA Driver API or statically-linked CUDA Runtime API - dynamically-linked CUDA Runtime API is not supported at all.
Common man. You can't call it a drop in replacement if it's an incomplete proof of concept.
I think calling it a drop in replacement highlights that the aim of Zluda is to not require recompilation of the software being run. Maybe “Proof of concept drop in replacement” would be a better description?
I was nearly put off by that but maybe it should be given a chance with some funding or what not. All great and interesting things come from proof of concepts, that actually work.
At least it isn't vapourware. Unlike some 'projects' I see on GitHub.
Literally the top comment is even suggesting that Intel could be interested in funding this, so surely it deserves a chance with some backing, even if it is 'incomplete'.
Care to explain the downvotes?
Any "professional" application solely use async APIs so while these numbers may look impressive something like tensorflow or pytorch would either not run or be incredibly slow.
Geekbench could, at best, be described as a fine benchmark. But it's pretty terrible in a bunch of respects.
I would love to see more benchmarks run through with this.
tl;dr: someone would need to re-implement cuDNN
Then how does InterlockedAdd HLSL intrinsic works on Intel? https://docs.microsoft.com/en-us/windows/win32/direct3dhlsl/...
There a few issues which come to mid though:
1. How do these benchmark results compare to writing x64_64 implementations directly? That is, is it enough to reach OpenCL-level performance?
2. What about if I want to use both a CPU and and a GPU? It seems that would not be supported, as this library produces a `libcuda.so`.
3. What about AMD chips?
4. The README says:
> Authors of CUDA benchmarks used CUDA functions atomicInc
> and atomicDec which have direct hardware support on NVIDIA
> cards, but no hardware support on Intel cards. They have
> to be emulated in software, which limits performance
is there really no support, not even in the near future, for atomic increment on Intel chips?
5. This is all written in Rust. It's a bit fishy to me for a C library to be implemented in Rust, though maybe I'm just a little prejudiced.
6. Where is the documentation of the semantic differences between proper CUDA and ZLUDA? e.g. - what do streams do? How do events work? etc.
If they did they would not have:
- Have had support for Linux only, with no ROCm on Windows
- Have made ROCm as a GPU-specific targeted process, there's no IR like PTX to make your current *binary* run on future GPU generations
- Having dropped support for GCN2/3 (https://github.com/RadeonOpenCompute/ROCm/issues/1353#issuec...) making the _only_ supported customer GPU generation Vega, with no support for RDNA/RDNA2.
They obviously don't care about the market as they should, despite anything they might or they might not say. Nothing to see here... It's only and solely their own fault that NVIDIA is the only option.
Intel so far has a much more competent strategy around GPU computing, and might prove to be an actual competitor. I've written off AMD as a possible competitor to NVIDIA for GPU computing a long time ago.
Intel not only has better OpenCL support but will come out out of the gate with OneAPI that will support all Intel GPUs this means productivity applications could use it wether it’s for a laptop or for a future productivity workstation with an Intel discrete GPU.
It’s pretty much impossible to buy a laptop with an AMD GPU and run ROCm on it, even the discrete cards are not officially supported and since the ROCm binaries are hardware specific without official support things tend to be even more broken than what they are now.
I really can’t understand how AMD could cock it up so badly.
The latest iGPU supports more than 1 TFLOP in GPU performance. This is apparently more than the performance of two years old Nvidia GeForce GT 1030.
[1]https://www.notebookcheck.net/Intel-s-Elkhart-Lake-SoC-will-...
Almost 4 years old...
https://www.techpowerup.com/gpu-specs/geforce-gt-1030.c2954
"The GeForce GT 1030 is an entry-level graphics card by NVIDIA, launched in May 2017."
Anyway having iGPU inside an entry level Intel's CPU with a potential of 32 GB of RAM is amazing. The performance is probably not that far off from my old high end Asus ROG gaming laptop that I bought in 2014 with Nvidia GTX 700 series GPU.
The laptop with the entry level iGPU will probably cost less than 20% of the original ROG laptop but now can supports up to three 4K monitors!
If Asus can resurrect their infamous Netbook series laptop or cheap laptop lines with these new entry level CPUs with iGPU I bet it can sell them like hot cakes.
[1]https://rog.asus.com/articles/g-series-gaming-laptops/asus-i...