Compiling Julia for NVIDIA GPUs
blog.maleadt.net
blog.maleadt.net
using CUDA
# define a kernel
@kernel function kernel_vadd(a, b, c)
i = blockId_x() + (threadId_x()-1) * numBlocks_x()
c[i] = a[i] + b[i]
end
# create some data
dims = (3, 4)
a = round(rand(Float32, dims) * 100)
b = round(rand(Float32, dims) * 100)
c = Array(Float32, dims)
# execute!
@cuda kernel_vadd(CuIn(a), CuIn(b), CuOut(c))
# verify
@show a+b == c
i.e. no set up, no "cuda context" (whatever that is), and no tear down afterwards. I understand that manual memory management is almost necessary with this application, but it seems that most of it could be automated in the most common case of "I have a couple of large arrays and a few operations I want to perform."Currently supports both including multithreading. Its just he unified interface that is in the works.
For single GPU stuff, I find myself writing the performance critical kernels in cuda, and using thrust for nearly all the easy stuff.
More involved example using a neural network: http://code.cogbits.com/wiki/doku.php?id=tutorial_cuda
https://github.com/takagi/cl-cuda/blob/master/examples/vecto...
I though that Julia had macros, so I don't understand why what you propose is not possible (note to self: find time to learn Julia).
When working in julia, what are the benefits of tying oneself to CUDA (and not running accelerated on on-die graphics or on amd gpus) -- or doesn't nvidia work reliably/well with opencl?
My project also provides compiler support for lowering Julia to CUDA assembly, so you don't need to write CUDA code yourself. Added to that, my runtime also contains (PoC) higher-level wrappers, making it easier to call CUDA kernels, upload data, etc.
Concerning tying yourself to the NVIDIA-stack: it's still the most mature and versatile toolchain, which is why I picked it in the first place. My long term plan was to switch over to SPIR (or some other cross-vendor stack) as soon as possible. At that point, switching user-code over to that new back-end would (theoretically) not require that much effort, since the kernels are written in julia-code instead of CUDA C (except for the runtime interactions, of course).
As for Nvidia/CUDA being more mature -- that was what I feared -- it seems a common sentiment in the discussions I've seen on OpenCL/CUDA.