Download Ollama on a modern MacBook and can run 13B and even higher (if your RAM allows) at fast speeds. People run smaller models locally on their phones
Google has trained their latest models on their own TPUs... not using Nvidia to my knowledge.
So, no, there are alternatives. CUDA has the largest mindshare on the training side though.
OpenMP is still a thing in 2024, but I presume that is not the kind of scale you are asking about.
If you have MI accelerators, you'll be using ROCm anyway (although AMD contracted Andrzej Janik in 2022 to make ZLUDA run on AMD GPUs. I have no idea what the practical applications of it are at the moment.)
The only other serious challenger (apart from the existing GPU manufactures like AMD) who is trying to give Nvidia a run for its money is Tenstorrent and their TT-Buda software kit.
[1@2024-04-06] https://www.youtube.com/watch?v=j7MRj4N2Cyk&t=429s
[Twitch] https://twitch.tv/georgehotz
Some quotes:
I find it incredible that these companies that have large support contracts with you and have invested hundreds of thousands of dollars into your products, have been forced to turn to me, a mostly unknown self-employed hacker with very limited resources to try to work around these bugs (design faults?) in your hardware.
In the VFIO space we no longer recommend AMD GPUs at all, in every instance where people ask for which GPU to use for their new build, the advise is to use NVidia.
[1]: https://www.reddit.com/r/Amd/comments/1bsjm5a/letter_to_amd_...
The guy did one good jailbreak for the iPhone, and as near as I can tell, the rest of his work has been a lot of boasting, half-assed hyped-up implementations (e.g: his self-driving car), and trying to befriend other powerful people in tech (see: his promise to single-handedly fix Musk's Twitter). He might be a smart dude, but he vastly overrates his own accomplishments, and doesn't finish near anything he starts.