There's also OpenAI Triton. People seem to miss that OpenAI is not using CUDA...
It also looks like they added MLIR backend to Triton though I wonder if Mojo has advantages since it was designed with MLIR in mind? https://github.com/openai/triton/pull/1004