Julia GPU
notamonadtutorial.com
notamonadtutorial.com
> Much of the initial work focused on developing tools that make it possible to write low-level code in Julia. For example, we developed the LLVM.jl package that gives us access to the LLVM APIs. Recently, our focus has shifted towards generalizing this functionality so that other GPU back-ends, like AMDGPU.jl or oneAPI.jl can benefit from developments to CUDA.jl. Vendor-neutral array operations, for examples, are now implemented in GPUArrays.jl whereas shared compiler functionality now lives in GPUCompiler.jl. That should make it possible to work on several GPU back-ends, even though most of them are maintained by only a single developer.
This is the takeaway to me. First-class access to GPU accelerators using the same syntax, regardless of the vendor :)
https://github.com/ROCm-Developer-Tools/HIP
I realize these sort of tools aren't magic and whatever it spites out will need work, but it seems like a really good thin starting place for AMD support with a lower overhead for growth.
After the original CUDA bits can ""cross-compile"", the workflow is greatly reduced, right?
Workflow:
- update CUDA code
- push through the HIPIFY tool
- Fix what is broken (if you can fix it on the CUDA side)
After enough iterations, the CUDA code will grow friendly to HIPification...
HIP and HIPify only work on C++ source code, via a Perl script. Since we start with plain Julia code, and we already have LLVM integrated into Julia's compiler, it's easiest to just change the LLVM "target" from Native to AMDGPU (or NVPTX in CUDA.jl's case) to get native machine code, while preserving Julia's semantics for the most part.
Also, interfacing to ROCR (AMD's implementation of the Heterogeneous System Architecture or HSA runtime) was really easy when I first started on this, and codegen through Julia's compiler and LLVM is trivial when you have CUDAnative.jl (CUDA.jl's predecessor) to look at :)
I should also mention that not everything that CUDA does maps well to AMD GPU; CUDA's streams are generally in-order (blocking), whereas AMD's queues are non-blocking unless barriers are scheduled. Also, things like hostcall (calling a CPU function from the GPU) doesn't have an obvious alternative with CUDA.
I think the relevant comparison today is with Numba, here's a real world recurrence analysis,
@cuda.jit
def _sseij(Y, I, J, O):
# strides
sty = cuda.blockDim.x
sbx = sty * cuda.blockDim.y
sby = sbx * cuda.gridDim.x
sbz = sby * cuda.gridDim.y
# this thread's index
t = (cuda.threadIdx.x
+ cuda.threadIdx.y * sty
+ cuda.blockIdx.x * sbx
+ cuda.blockIdx.y * sby
+ cuda.blockIdx.z * sbz)
if t < I.size:
i = I[t]
j = J[t]
x = nb.float32(0.0)
for k in range(Y.shape[0]):
x += (Y[k, i] - Y[k, j])**2
O[t] = math.sqrt(x)
most of it is index calculation, but super easyC isn't usually a breath of fresh air, but the last two times I tried to use nubacuda and failed (~1.5 year ago), it sure felt that way.
EDIT: yes, looks like it supports on-device buffers now. I'm still in "once bitten, twice shy" mode on account of the bugs and debug story, but I'm cautiously optimistic.
"You're a CUDA programmer. Well, why don't you learn it all over again the Julia way, and also unlearn CUDA, so the next time you have to program an nVidia GPU, you don't have a choice but to do it in Julia"
Vendor lock-in FTW.
Out of the frying pan, into the fire. Great.
When any code is written in some language?