I am confused. Is it another competitor of Tensorflow, JAX, and Pytorch? Or something else?
I imagine these libraries and possibly some users would implement libraries on top of this language and reap some of the optimization benefit without having to maintain low-level CUDA specific code.
obligatory reference to the family of work: https://github.com/merrymercy/awesome-tensor-compilers
Edit: should've read the post before commenting. Looks like they are in fact using LLVM's PTX backend (ie generating cuda kernels from scratch). Kudos to them