It was Nvidia’s competitors’ job to ensure this never happened but hardware companies rarely value the software stack as much as they should have. AMD screwed up several times to build a similar tool and Intel did not manage to create a proper programming model for their vector instruction sets despite having some promising internal efforts. Cuda allowed AlexNet to usher the deep learning era, written by university researchers. The flash attention implementation allowed quadratic attention to have reasonable memory requirements again written by a university professor and used by everyone. If I were to complain about Nvidia, I would complain about them abandoning the personal computing track for datacenter profits, definitely not about pushing Cuda.
Translating CUDA kernels to ROCm (or any other ecosystem, tbh) is going to become trivially easy with AI. Not saying that code is the only glue that Nvidia has when it comes to their ecosystem, but it’s certainly a large part of it.
It's not that difficult at the moment. Don't know why you would need an LLM to translate between two nearly identical language extensions.
Optimization is different, but that's not really a translation level task.
If an LLM can do a rote task for me, I don’t see why I wouldn’t use it
Because it's more expensive and more likely to fuck it up.