Vortex: OpenCL compatible RISC-V GPGPU
vortex.cc.gatech.edu
vortex.cc.gatech.edu
i imagine if i tried to explain to someone the state of modern GPU computing i'd sound like an insane person, why is it so complicated, ugh :(
so there's openGL and openCL, which are specifications with various implementations, both widely supported, but not as performant/modern
there are vulkan/metal/dx, which are all modern with good driver support (metal limited driver support) graphics-first APIs, with compute shaders (vulkan kompute, metal performance shaders and whatever dx has?), all three have excellent performance
then there is cuda/hip, both of which are proprietary, cuda w/ excellent driver support, and hip w/ reportedly awful driver support, cuda can only target nvidia gpus, and hip can target both cuda and amd gpus (though very few amd gpus), both have very good performance
then there is SYCL, which is by khronos, and it's what intel uses w/ openAPI, which is a heterogenous computing framework that can produce openCL/vulkan and in the future other api's code
then there's webgpu, which is a spec, that has two C and one rust implementations, that can target all of the above (except SYCL obv, bc it's not a graphics api)
did i get that right?
I mean, do you think this is different for anything else in our field? This was in some sense really simple... try doing this for front-end GUI frameworks or even multi-tenant isolation technologies.
We're seeing a version of that replaying, with some parts better and some parts worse.
The GPU situation would be ripe for regulation to lift GPU software development from the dark ages, but companies who benefit most (NVidia) have probably become too big to regulate.
I guess you are from the US and not EU? It's when they get too big and abuse their monopoly they will be regulated.
AFAIK nvidia have not done anything to abuse its power yet, but its hard for big companies to not use their monopoly in a illegal way since its always a good a Market strategy just illegal
To be fair there was some kind of support on the PS3 with GL ES 1.0/Cg shaders as secondary API, later dropped, Wii GX(2) shaders are based on GLSL, and while Switch does support GL 4.6 and Vulkan for porting purposes, it is either Nintendo extensions spaghetti time, or rewrite in NVN, for actually taking full advantage of the hardware.
This is explains the rise of ROCm and DPC++ as systems having equivalent higher level constructs to Cuda.
However almost no one cared, meaning AMD and Intel never shiped anything worth using, so OpenCL 3.0 is OpenCL 1.0 rebranded, and out of it, the C++ efforts were placed into SYSCL intead.
Which Intel picked up for their own DPC++ efforts, an additional tooling layer on top of SYSCL, meanwhile the only company selling usable SYCL developer experience is a former compiler vendor, that used to work with Sony in high performance compilers for the Playstation (Codeplay), as they pivoted away from console development.
Eventually Intel acquired Codeplay, and they are now the main supporters of One API tools, and the whole UXL efforts.
In the middle of all this, AMD decided to go with their own efforts.
It isn't only up to NVidia, when their competition can't get their story straight for decades.
If anything the situation is that Nvidia is far-sighted, developing hardware and software for general purpose GPU computing and the other chip manufacturers still think like traditional (non-CPU) chip manufacturers - they just want to sell a best-chip for a single purpose. This explain why they throw random things against the wall to see if they stick rather than choosing a general purpose and committing to it indefinitely.
People have even started using Vulkan in its place for some GPGPU, because at least Vulkan has good drivers everywhere, being a low-level API.
The problem is that the drivers are merely OK. Presumably if you're using OpenCL you care about the performance (otherwise why would you??) and since that's the case, it's the best on no platforms, and there are alternatives for any set of platforms that do better.
I think OpenCL is sadly on its way out, and it's mostly Apple's fault (and Nvidia a little). Vulkan compute is much more interesting if you're looking to leverage iGPUs/mobile/other random CPUs.
If you're targetting workstations/server workloads only, it makes sense to restrict yourself to a subset of accelerator types and code for that (eg. Torch or JAX for GPUs, use highway for SIMD, etc.)
I'm definitely not saying OpenCL is any sort of a reasonable default for cross platform GPGPU work. In truth, I don't think there is any reasonable "general" default for that sort of thing. Vulkan has its own issues (only works via a compatibility layer on MacOS, implementation quality varies widely, extension hell, boilerplate hell, some low level things are just impossible, etc.) and everything else is a higher level approach that can't work for everything by definition.
It's a pretty sad situation overall and every solution has severe tradeoffs. Personally, I just write CUDA when I can get away with it and try to stick to OpenCL otherwise, but everyone needs to make that choice for their own set of tradeoffs.
But the driver implementations inconsistency, version support issues, etc. meant people used CUDA instead.
I agree Vulkan has its own issues, and having written some MoltenVK stuff, you clearly know the quality-of-life pains in developping with it. That said, at least from the user side it works and performs well.
That being said though, for how little funding, research, and support OpenCL receives, it is an astoundingly capable tool.
I don't think it's fair to compare apples and oranges, even if they're clearly both "fruit" with a "high sugar content".
Apple gave up on OpenCL, after disagreements how Khronos managed it.
Metal Compute provides a CUDA like development experience.
Ah yes, the old "I can't completely control it, cause I'm Apple, so I'm taking my ball and leaving" ploy.
Though, I will admit, AMD was also playing this game at the same time, so a disagreement was bound to happen between them and Apple.
> Metal Compute provides a CUDA like development experience.
And there it is. Proving my prior statement, Apple didn't want an open compute platform, what they wanted was their own compute platform and it would have "been nice" if they could pretend it was "open" and they "shared" nicely with the other tech kids.
Before people come at me as an Apple hater, please know that there are many, many Apple devices in my home. I'm not an Apple hater, but I call it like I see it.
And VS integration means intellisense and syntax highlighting as well, not just calling their compiler.
[PDF] https://john.cs.olemiss.edu/heroes/papers/AMD_OpenCL_Program... See section 3.1 for debugging, section 4.2.2 for kernel code editor.
With a parallel CPU implementation I have pretty few debugging issues, and most importantly, I definitely enjoy having excellent performance on Intel, AMD and Nvidia GPUs, on Windows, Mac and Linux.
Being able to declare variables anywhere can lead to some really nasty bugs especially when using goto.
It would make for a helluva blog post, though, or set up a shopify selling pre-made ones for people like you :-)