Do you have another open, cross-platform, widely compatible GPU programming framework to recommend?
On iOS it is kind of deprecated and the way forward is Metal Compute.
On Android it never happened, Google created their own Renderscript dialect instead.
The alternatives recommended here aren't even serious IMHO. I'd rather switch to CUDA and wait till Intel/AMD sort out a REAL compatibility layer than deal with those.
Unless I'm mistaken, HIP still requires a separate compile for either platform and what runtime do they expect end users to have exactly?! At least CUDA and OpenCL are integrated in the vendor drivers.
Vulkan compute with SPIR-V seems to be the only real solution, but even that is still very early. Sill waiting for proper OpenCL 2.0 support in NVIDIA drivers :P
(The models can be executed on low-powered, commodity CPUs. No need for any GPU there.)
That's totally an option for our product, great idea! Why did I never think of this!
No seriously we are shipping, using OpenCL and it gives about a 20 times performance advantage for most users regardless if they have AMD or NVIDIA hardware. If something that's actually better than OpenCL comes along (or if AMD RTG goes out of business) I'll switch to it no heart broken.
But that hasn't happened yet.
Do you mean adaptive algorithms or dynamic recompilation? And yes I do expect that cross API will be difficult, for both running and getting good perf.
But it is not just the room at the high performance end of the spectrum that is important, but also the lower end that is stifled by the barrier to entry that would benefit from the extra compute.
My point is - was drawn to OpenCL for its 'portability' claim, and yes kernels will 'run' on various hardware, but with massive differences in speed. what good does that portability do me? My workloads (scientific computing, branch heavy) are different from the typical ML half- or single precision MulDiv()/linear algebra applications, so all the hand-tuned CUDA libraries aren't even my concern. The elephant in the room is that performance doesn't depend on this API or that; it's in how you tune your algorithm to the actual hardware you're running on. Which is the direct opposite of portability.
To come back to the post I was replying to - yes it's 'trivial' (I mean, a lot of work, but technically not hard, not to belittle your work) to compile almost any statically typed code into either SPIR-V or PTX or any other future format for that matter. But that doesn't mean that it will work 'well' (not even 'optimal') on other hardware. In the real world, you're almost always better off just spending a few hundred to get the same GPU as whoever wrote the kernels tested them on (or if you're running your kernels on existing clusters, to focus on optimizing them for what you know you'll be running them on).
Oh and all of that is just considering GPU's. I mean, when you read an OpenCL book they make it seem (in the first few chapters) that you don't even have to think about whether you'll be running on a CPU or a GPU. And then you accidentally use USE_HOST_PTR instead of COPY_HOST_PTR (or the other way around) in the wrong place, and all of a sudden your code is 10x slower than the sequential version of your algo even.
What I'm saying is - I no longer believe in easy to use abstractions for these purposes. If you want speed, you code to the metal, and/or you tweak your abstractions to your specific use case. Yes this is a lot of work. And if you don't need speed, you just throw in a few std::thread's here and there and call it a day.
But that's just my conclusion for my use cases.
I agree that it is not easy to 'parameterise the metal', but it is certainly doable in D[1], in C++ (guessed from the std::thread) good lord no: D's is orders of magnitude ahead of C+'s meta-programming. Writing different versions of a kernel is a poor mans parameterisation ;)
LDC, the LLVM D Compiler, will be getting a dynamic re-optimisation, which I hope to get to play nicely with DCompute. Well PTX, because SPIR-V doesn't yet have a jit backend. Dynamic re-optimisation from what I understand is freeze some variables as constants and rerun the optimisation passes. This is as opposed to recompilation with complete restructuring of the kernel. Not quite the same but for things like loop counts and branch "prediction" this should help a lot.
w.r.t USE_HOST_PTR/COPY_HOST_PTR, that's not a part of the kernel parameterisation, that's part of the host and is easily adjustable. Yes you need to figure our which one to use, but that's just part of tuning.
> What I'm saying is - I no longer believe in easy to use abstractions for these purposes.
I hope to be able to show you otherwise :)
[1](https://github.com/libmir/mir-glas#porting-to-a-new-target)
If AMD GPUs die out, CUDA it is.
Or maybe give up and use a wrapper library for whatever 10 alternatives-only-supported-by-one-marginal-vendor there are. (like Apple)
The issues go away if you use a good OpenCL frontend. PyOpenCL for Python goes a long way towards this, and is not really any more awkward than the corresponding PyCUDA, and higher-level languages that generate OpenCL code, like Lift[0] or Futhark[1] (tooting my own horn here), remove the awkwardness completely.
[0]: http://www.lift-project.org/ [1]: https://futhark-lang.org
OpenGL ES only took off thanks to gaming on the iPhone, and now is deprecated on Apple platforms.
Vulkan still lives in a C world, and the semi-official C++ bindings only exist thanks NVidia.
OpenCL waited too long to support C++, Fortran and providing an infrastructure for compiler writers to add GPU support to their own languages. And two years later the majority of drivers are not there yet.
Which Khronos finally adopted as SPIR, but the drivers aren't there yet.
Regarding the other Khronos APIs, a C API is like being stuck in a PDP-11 world.
Many mix the idea of C API with OS ABI, it only happens to be the same if the OS APIs are exposed as plain C.
There many cases where this isn't like it, e.g. mainframes, mobile OSes, and most userspace on OS X (Obj-C runtime) and Windows (.NET and COM).
Yes driver support is a bit lacking, although I hop that I can convince the OpenCL working group of the need to get a backend (such as https://github.com/thewilsonator/llvm-target-spirv) into mainline LLVM so that writing drivers becomes easier for vendors.
And segmenting the codebase is a MAJOR feature. With CUDA you are stuck on an old compiler until NVIDIA issues an update. How anyone can think this is a good idea...
"Designing (New) C++ Hardware”
https://www.youtube.com/watch?v=86seb-iZCnI
CUDA has had C++, Fortran support since the early days, with PTX for additional compiler backends added in version 1.4.
That was 2007, Khronos waited until 2015 to specific similar capabilities.
People still write regular shaders in languages that are much more C like.
People writing GPGPU code are few. Most of the DL GPU use is in Python through several layers, and in the end you are running hand-written SASS assembly sitting in an NVIDIA DLL or whatever.
I guess some people must think it's handy to have C++ support in GPU kernels, or they wouldn't have added the feature. But for it to drive technology, hard to believe.
The Metal, DirectX, PS3, PS4, Nintendo and several middleware engines are C++ like.
Also the fact that OpenCL lost to CUDA for being stuck in C for so long, shows what most GPU devs actually prefer.
Aw come on, you're sure it has nothing to do with the largest GPGPU vendor pushing CUDA very heavily and intentionally gutting their OpenCL tools? Or putting out a ton of very high performance libraries with no OpenCL equivalent? Putting out a shitton of marketing and tutorial videos for CUDA only?
Yeah, that surely was totally unrelated.
Also pushing a proprietary standard goes faster than a standardized one. No surprise there.
If AMD, Intel and embedded OEMs actually produced quality OpenCL drivers, debugger and IDE support and libraries that could match CUDA productivity, maybe devs would bother to use C with OpenCL.
Even Google decided to create their Renderscript dialect instead of supporting OpenCL on Android.
I am saying that if the other GPU vendors bothered to actually provide a competing technology stack, that was worth the pain of using plain C, maybe GPU devs would have bothered.
You're not wrong there. However with the advent of SPIR-V it is possible to write code in whatever language you please (with the caveat that at the moment you need an LLVM backend using https://github.com/thewilsonator/llvm-target-spirv or the Khronos repo I forked that off of. Then comes the issue of making the code generator friendly interface user friendly, which I have done for D so that you get the ease of use of CUDA.