Interactive GPU Programming, Part 1: Hello CUDA
dragan.rocks
dragan.rocks
If you already know Clojure, this is probably the best chance to extend something you already love using. If you don't, you're probably better off learning either CUDA or Clojure rather than both at the same time. Debugging CUDA errors are already painful, I wouldn't add a new host language on top of that.
For context, I'm currently taking my school's GPGPU course. We've just started actually writing non-trivial code.
Looks like Khronos are looking into converging them in some way: https://www.pcper.com/reviews/Graphics-Cards/Follow-Neil-Tre...
CUDA unfortuantely is Nvidia's lock-in, so not a good way forward.
While I agree with your sentiment, unfortunately Nvidia is the only vendor that pays considerable number of people to develop the ecosystem. AMD basically says "get lost" by refusing to put more than a handful of people on the job of providing OpenCL libraries. And, BTW, they change their minds every few years. I hope that HIP won't be abandonware...
Vulkan itself is developed and supported well, and it already can be used for compute as far as I know. But apparently there are some features that come from the OpenCL world that need to be filled in. It wouldn't be AMD's exclusive effort. So hopefully things will start moving.
It took them being beaten by NVidia to actually care to add SPIR and C++ support to OpenCL.
Even now, while CUDA brings C++ compiler out of the box with their SDK, for OpenCL one needs to go to Codeplay and download their ComputeCpp Community edition compiler for SYSCL support, that might or not, support a given card. Hardly any better.
E.g. ccminer. Try to make it, but it finds my modern gcc or clang too modern :(
This requires nvcc and the device compiler to have exact knowledge of how the host compiler compiles every single construct (thing e.g. about alignment and padding in complex structures), and they must at least be able to parse the syntax of the host include files (which e.g. fails if the include files have C++11 syntax, but the device compiler only knows how to parse C++98).
http://nd4j.org/ - in built GPU garbage collector and everything.
If you want raw cuda primitives (not generally recommendended and hard to do right) - you can take a look at our javacpp based (we also maintain this) cuda bindings: https://github.com/bytedeco/javacpp-presets/tree/master/cuda
Unlike jcuda (which people typically recommend despite not being updated as often) we actually depend on this for the nd4j and deeplearning4j projects.
These cuda bindings are meant to be a 1 to 1 mapping to the cuda api as well. Hope this helps!
If you want a fairly small and minimalistic look at the underlying c code which uses cuda take a look at: https://github.com/deeplearning4j/libnd4j
All of this is published on maven central for you and runs on linux, windows and even mac. It's also the same api. All you do is switch the backend.