nVidia’s top management did amazing job with their strategy. They have spent decades, and millions of dollars, developing and promoting CUDA. Now this motivates people to spend that amount of money on nVidia’s hardware.
I don’t believe the trend continues for ML inference, at least not for long. Unlike traditional HPC stuff like FEM or fluid dynamics, the GPU code of ML inference is not that complicated. The complexity is limited to the data being processed, these ML models.
It’s not terribly hard to port the inference to other hardware architecture or software stack, reusing the models. I think it’s only a matter of time until AI companies like OpenAI will start porting their code from CUDA to VulkanCompute or some other vendor-agnostic stack, to avoid buying or renting such $100k computers. Don’t get me wrong, these computers are awesome, just very expensive.