Looks like this entire paper is just about how to move/remove these barriers.
Looks like this entire paper is just about how to move/remove these barriers.
An interesting add on to me would be the handling of conditionals. Because newer GPUs have independent thread scheduling which is not present in the older ones, you have to wonder what is the desired behaviour if you are using CPU execution as a debugger of sorts(or are just GPU poor). It'd be super cool to expose those semantics as a compiler flag for your transpiler, allowing me to potentially debug some code as if it ran on an ancient GPU like a K80 for some fast local debugging.
But the ambitious question here is this - if you take existing GPU code, run it through a transpiler and generate better code than handwritten OpenMP, do you need to maintain an OpenMP backend for the CPU in the first place? It'd be better to express everything in a more richer parallel model with support for nested synchronization right? And let the compiler handle the job of inter-converting between parallelism models. It's like saying if Pytorch 2.0 generates good Triton code, we could just transpile that to CPUs and get rid of the CPU backend. (of course triton doesn't support all patterns so you would fall back to aten, and this kind of goes for a toss)
I agree that statically proving that something like the syncing is unnecessary can only be a good thing.
The question of why not simply take your GPU code and transpile to CPU code is more of the question of what did you originally lose in writing the GPU code to begin with. If you are talking about ML work most of that is expressed a bunch of matrix operations that naturally translate to GPUs with low impedance. But other kinds of operations might be better expressed directly as CPU code (any serial operations). And for CPU to GPU the loss as you have pointed out is probably in the synchronization.