Does this solve some of the core problems with GPU programming, such the difficulty of writing reusable code without significant performance overhead (via nested parallelism or fusion), or the need to follow some potentially awkward rules for performance reasons (struct-of-arrays and coalesced memory access patterns?).
I mean, it is surely nice to easily launch a bunch of threads within a single-source program, but there are already plenty of C++ libraries that let you do this, and it has not really led to an explosion of making efficient GPU programming accessible to the layman.