We recently showed DiffEqGPU.jl generating customized ODE solver kernels for NVIDIA CUDA, AMD GPUs, Intel OneAPI, and Apple Metal, where for CUDA it matches state of the art (MPGOS) which is about 10x-100x something like Jax/PyTorch (where the performance difference comes from inefficiencies of using vmap vs actually writing and calling a kernel). It's all in https://arxiv.org/abs/2304.06835. So this stuff exists and people are using it. Of course the caveat here is this is the context of engineering applications so someone would need to do similar for LLMs to fully relate back to the article, but it shows the tools are ready to a large extent for someone to step up in the ML space.