I have heard about Go's parallelism being awesome, does anyone know of a comparison between Go's and Julia's models?
I have heard about Go's parallelism being awesome, does anyone know of a comparison between Go's and Julia's models?
TLDR: The entire GPUArrays.jl test suite now passes with AMDGPU.jl. There are still some missing features and it is not as mature as the nvidia version, but this space is progressing rapidly, and benefited from the generic GPU compilation pipeline that was initially built for CUDA.jl
The problem with AMD GPGPU is not software, it is that AMD literally does not care.
Rocm gets releases every few months or so. The llvm project part is mirrored to GitHub in real time.
I think AMD has some work to do on non-C++/Python ecosystem engagement for sure, but they've built a foundation that's quite easy to build upon and get excellent performance and functionality; AMDGPU.jl is a testament to that.
There's nothing fundamentally special to Go's "parallelism".
It uses near-preempted (and now actually preempted I think, since a few versions back? before that they'd have implicit yield points at function calls but without function calls you could lock out the scheduler) userland threads with small stacks.
The parallelism comes from the m:n scheduling. Go has no constructs for massive parallelisation, where you hand it a loop and it efficiently runs that on a hundred cores, at least not built in.
Reasoning about threads & basic blocking calls is a simpler mental model compared to async/await style concurrency in my opinion.
Go is fundamentally designed for developing backend applications, not high perf mathematical/scientific computing. Features like unrolling for-loops across hundreds of cores or offloading to GPU's would be out of place in the language.
Well... there's still a problem in the tens-of-millions of Goroutines.
But the issue is that OS-threads generally had issues at the ~100,000 of OS-threads. Making something "more lightweight" than pthreads makes sense, because 10,000,000 coroutines is a fundamentally different program design than 100,000 pthreads.
With threads I find I can more easily mentally organise which code in my application is executing in parallel.
In async/await land, you end up with 2 classes of functions async & non async functions, it's up to the callee to determine wether it is a blocking call or not. With threads you just block by default, and let the caller determine whether it's appropriate to block or execute the function on another thread. See https://journal.stuffwithstuff.com/2015/02/01/what-color-is-... for someone mere eloquent than myself.