Since I only have one GPU usually, I'm still waiting for the day I can do this:
using CUDA
# define a kernel
@kernel function kernel_vadd(a, b, c)
i = blockId_x() + (threadId_x()-1) * numBlocks_x()
c[i] = a[i] + b[i]
end
# create some data
dims = (3, 4)
a = round(rand(Float32, dims) * 100)
b = round(rand(Float32, dims) * 100)
c = Array(Float32, dims)
# execute!
@cuda kernel_vadd(CuIn(a), CuIn(b), CuOut(c))
# verify
@show a+b == c
i.e. no set up, no "cuda context" (whatever that is), and no tear down afterwards. I understand that manual memory management is almost necessary with this application, but it seems that most of it could be automated in the most common case of "I have a couple of large arrays and a few operations I want to perform."