Optimizing a Rust GPU matmul kernel
rust-gpu.github.io
rust-gpu.github.io
Any pointers to examples? I'd be fine sticking with floats as a first step, but would love to see some (reasonably optimized) low-level GPU code for working with matrices of complex numbers (preferably for Metal or WGPU, which is what I'm using).
For Rust GPU, nothing built in but there are libraries like https://github.com/rust-num/num-complex that support `no_std` and should work on the GPU. I've never used them so I don't know what (if any) the perf hit would be.
The main issue is that double precision is not so interesting for AI and graphics, and so silicon is rather spent on more of these features and less double precision. Not so for HPC, though, and GPUs specialized for this usually have better throughput. For example, the AMD MI210 has the same performance in single and double precision (matrix) operations, while graphics GPUs either have something like 1/2, 1/4, 1/16 etc rate of fp64:fp16, or have no support at all.
like tensor core support, non-uniform thread groups, different data type support, a bunch of simd group functions, stuff like that…
i couldn’t find info wrt this in rust gpu, so i’m assuming it just tries to target the narrowest available feature set, that’s compatible across all shading languages?
> These Rust GPU programs are then compiled into SPIR-V, a low-level format that most GPUs understand
(Culturally of course the big one is the fragmentation and the proprietary nature of everything which is the reason so little gets done on GPUs and the horror of attempting multiplatform software there)
Correct me if I'm wrong or misunderstanding you, but doesn't SPIR-V support tensor cores via SPV_KHR_cooperative_matrix?
At first blush it just sounds like something to allow multiple shader compute elements to work more efficiently together ("cooperate") on a single bigger matrix computation.
> Additionally, if the GPU includes dedicated hardware for high-speed matrix operations, such as the Tensor Cores on Turing GPUs, then the Cooperative Matrix extension can tap into the power of this acceleration with no application changes.
The benchmark graph doesn't look too great, though - around half the "theoretical peak tensor core performance".
[0]: https://developer.nvidia.com/blog/machine-learning-accelerat...
We already have a lot (we have many intrinsics and an `arch` module like `std::arch`, asm! to include raw spirv, etc). For example, here are the intrinsics: https://rust-gpu.github.io/rust-gpu/api/spirv_std/arch/index... and here is support for ray tracing (which obviously is not on every card: https://rust-gpu.github.io/rust-gpu/api/spirv_std/ray_tracin...).
Vulkan has a way to query and specify different GPU capabilities and Rust-GPU uses that.
Rust and Vulkan have many of the tools we need for progressive enhancement, we are not focused on lowest common denominator.
A great technique is called 'tagless final encoding' [1]. Using this technique, you can specify capabilities of an embedded domain-specific language (eDSL) such that you can have a shared (but narrow) common set of features, while allowing specializations of this eDSL to support extra features.