Related question, what is the best way to handle kernel compatibility for CUDA, OpenCL, etc ... ?
I had to write a cross-platform kernel a few weeks ago, and I ended using pre-processor guards to make it work with the OpenCL and CUDA compilers [1].
[1] https://github.com/RaphaelJ/libhum/blob/main/libhum/match.ke...