> Can those command graphs loop?
Not directly (barrier packets wait for 0 only, plus the queue packets aren't preserved), but kernels can write to any dispatch queue themselves, so you can get the same effect at the end of your "loop body".
> I know this interface is undocumented... but I've had an idea for a GPU-language akin to Java or Lisp memory-management. The gist is that kernel_X() can execute, but may fail in any new() or malloc() command.
Memory allocation isn't special and allocators can be layered: you can allocate memory ahead of time and then just run the allocation algorithms GPU side. I wrote a Rust framework which cross compiles code/MIR on demand; you can in theory have a Rust allocator and use it to allocate GPU/CPU memory from either GPU/CPU. The only part the GPU can't do (directly) is invoke syscalls, which you can probably guess is the part needed to allocate virtual memory from the OS.
But as long as your allocator has enough spare virtual memory, it shouldn't need to do a syscall. And if you /really/ needed the ability GPU side, technically with signals you can actually just ask the CPU to allocate the virtual memory on the GPU's behalf and have the GPU spin until the allocation is "complete". Or with compiler support: automatically make the workgroup/kernel async and resume execution by enqueuing another kernel, but that sort of thing is kinda hard :).
Btw, the Rust framework is here: https://github.com/geobacter-rs/geobacter. I mostly work on it in my spare time, a scarce resource these days, so I admit it has some scuff.
> In such a case, I'd want the compute-graph to loop: while (kernel_X fails due to out-of-memory){ garbage_collect(); try kernel_X() again}.
Pretty much. Even garbage collection can (theoretically lol) happen on the GPU.