[1] https://gerritnowald.wordpress.com/2022/10/03/simulating-rot...
Our Julia code is parallelised with FLoops.jl, but so far Numba has shown surprising performance benefits when executing code in parallel, despite being slower when executed sequentially. Therefore I can imagine that Julia might yield better results when run in a regular desktop environment.
To bring Julia performance on par with the compiled languages I had to do a little bit of profiling and tweaking using @views.
https://github.com/JuliaParallel/rodinia/tree/master/julia_m...
It was touched 9 years ago, but maybe you have ported it to current standards. I don't think we had multithreading at that time, only multiprocessing.
Is your Julia implementations available somewhere? (Sorry if it is in your paper but I missed it). I vaguely remembered in the past that working with threads leaded to some additional allocations (compared to the serial code). Maybe this is also biting us here?
As far as I know the code was ported to use @floops, with minor optimisations in addition to that.
I think it's quite possible that it's an allocation issue, that's something we're looking into, although I don't have any specific results for Julia yet.
I am not familiar with Julia nor Numba internals, but maybe Numba, due to being more specialized, can actually provide LLVM with info that allows it to make more aggressive optimizations more easily.