Kompute – Vulkan Alternative to CUDA
github.com
github.com
But some of the caveats for compute applications are currently:
- No bfloat16 in shaders
- No shader work graphs (GPU-driven shader control flow)
- No inline PTX (inline GCN/RDNA/GEN is available)
These may or may not be important to you. Vulkan recently gained an ability to seamlessly dispatch CUDA kernels if you need these in some places, but there aren't currently similar Vulkan extensions for HIP.
i would imagine nvidia to not want their moat with CUDA be uprooted, and so their vulkan support might get gimped in some ways (that does not affect games). At least, that is what the cynic in me would predict.
AMD, very recently have also been putting a bit more work into it, though their drivers are still the worst, and still way worse than their pre-rocm drivers
Also one of my favourite non-VFX projects: 'Gollygang Ready' uses it to accelerate reaction-diffusion simulations.
I (perhaps naively) thought that OpenCL would still be a thing in a post-vulcan world.. SideFX are working on a new viewport using it.. I guess they solve different problems so can still coexist?
If I understand it correctly, that means the compute code is hardcoded into a specific assembly for gpu and to make it work with another card or newer one, you have recompile.
Like... Why? What is the problem with using SPIR-V, PTX or plain LLVM IR.
If we lived in a monoculture (e.g x86/64 for desktop apps), it would make sense, but there is a plethora of options and one gpu is not assembly compatible next gen.
Of course, using PTX does not necessarily guarantee backward compatibility, because you may use some PTX features that are supported only on newer NVIDIA GPUs. Nevertheless, inline PTX should continue to work on future GPUs.
Perhaps by "hardcoded" you have referred to only a part of your quotation, i.e. to "(inline GCN/RDNA/GEN is available)".
In this case I agree with you, but even so, there are enough cases when it is impossible, at least with the current compilers, to obtain the maximum performance allowed by the hardware without using inline assembly language, either for GPUs or for CPUs. Therefore it is good for the high-level programming language to permit the use of inline assembly language, even if this facility should not be abused.
Vulkan does make it easy to check which the feature availability of your device, and dispatch different shaders accordingly.
That said I've actually never had a reason to use inline assembly in shaders, personally.
It may, _in principle_, have been developed - with much more work than has gone into it - into such an alternative; but I am actually not sure of that since I have poor command of Vulcan. I got suspicious being someone who maintains C++ API wrappers for CUDA myself [2], and know that just doing that is a lot more code and a lot more work.
[1] - I assume it is opinionated to cater to CNN simulation for large language models, and basically not much more.
Does anyone have real world experience using Vulkan compute shaders versus, say, OpenCL? Does Kompute make things as straightforward as it seems?
The downside of compute shaders, regardless of graphics API (gl, dx, vk), is that they have an unspecified time limit after which the OS will kill your process. There are ways to disable this in your OS/GPU configuration but there isn't a portable way of doing this programmatically from your code.
Another issue is that if you use the same GPU for compute (or heavy graphics tasks) as your display output, your desktop responsiveness may go down. I had some issue where my desktop got really sluggish when I was drawing some graphics that took 350ms per frame (laptop integrated graphics).
SYCL is the unofficial successor to OpenCL - in that SYCL implementations like OpenCL are based on SPIR-V compute 'kernels'. (Note that these are not directly compatible with SPIR-V compute 'shaders' as found in Vulkan, so implementing OpenCL or SYCL on top of the Vulkan compute facilities comes with some challenges.)
Until then they are wannabe alternatives, for a subset of use cases, with lesser tooling.
It always feels like those proposing CUDA alternatives don't understand what they are trying to replace, and that is already the first error.
The product comes before the community (unless you have insane marketing money)
Alternatives are supposed to cover all uses cases, otherwise they aren't alternatives.
Not even AMD and Intel are able to make it happen, so it remains to be seen how much small communities are able to achieve.
A bicycle is an alternative to a car, however it doesn’t cover all the same use cases
I can't think of a single tech product where feature parity was a driver for growth. If all you have is parity, then all you compete on is lower price.
Usually, some advantage (better/safer/faster/easier) makes a difference for a few important use cases (good examples were early, feature-incomplete no-sql databases that excelled in one use case that existing SQL servers did not). That advantage hasn't emerged yet, so no community has developed.
We'll see if it ever does...
The fact many think CUDA is C/C++, is already the first error trying to replace it.
Wgpu is a bit behind on GPU features, for example they've added support for shader subgroups (aka warps or waves) in 2024, where as this feature was available in Vulkan 1.1 released in 2018 or Direct3d 12 shader model 6.0 (same timeframe). Wgpu still does not support buffer device address ("GPU pointers") which I consider quite a game changer.
Many popular tools, e.g. RenderDoc, don't have wgpu support.
If you are not targeting the web platform, wgpu isn't really bringing anything to the table.
If you are interested to learn more, do join the community through our discord here: https://discord.gg/MaH5Jv5zwv
For some background, this project started after seeing various renowned machine learning frameworks like Pytorch and Tensorflow integrating Vulkan as a backend. The Vulkan SDK offers a great low level interface that enables for highly specialized optimizations - however it comes at a cost of highly verbose code which requires 800-2000 lines of code to even begin writing application code. This has resulted in each of these projects having to implement the same baseline to abstract the non-compute related features of the Vulkan SDK.
This large amount of non-standardised boiler-plate can result in limited knowledge transfer, higher chance of unique framework implementation bugs being introduced, etc. We are aiming to address this with Kompute. As of today, we are now part of the Linux Foundation, and slowly contributing to the cross-vendor GPGPU revolution.
Some of the key features / highlights of Kompute:
* C++ SDK with Flexible Python Package * BYOV: Bring-your-own-Vulkan design to play nice with existing Vulkan applications * Asynchronous & parallel processing support through GPU family queues * Explicit relationships for GPU and host memory ownership and memory management: https://kompute.cc/overview/memory-management.html * Robust codebase with 90% unit test code coverage: https://kompute.cc/codecov/ * Mobile enabled via Android NDK across several architectures
Relevant blog posts:
Machine Learning: https://towardsdatascience.com/machine-learning-and-data-pro...
Mobile development: https://towardsdatascience.com/gpu-accelerated-machine-learn...
Game development (we need to update to Godot4): https://towardsdatascience.com/supercharging-game-developmen...
That first generation also included the ability to apply a handful of fixed activation functions, but really that's about it. The array is bigger than 32x32, also.
This is a big part of why RISC didn't win and today the largest server chips in use in datacenters are still mostly compatible with an 8 bit part from the early 1980's.