Comparatively few people have “deep” experience with CUDA (basically Tensorflow/Pytorch maintainers, some of whom are NVIDIA employees, and some working in HPC/supercomputing).
CUDA is indeed sticky, but the reason is probably because CUDA is supported on basically every NVIDIA GPU, whereas AMD’s ROCm was until recently limited to CDNA (datacenter) cards, so you couldn’t run it on your local AMD card. Intel is trying the same strategy with oneAPI, but since no one has managed to see a Habana Gaudi card (let alone a Gaudi2), they’re totally out of the running for now.
Separately, CUDA comes with many necessary extensions like cuSparse, cuDNN, etc. Those exist in other frameworks but there’s no comparison readily available, so no one is going to buy an AMD CDNA card.
AMD and Intel need to publish a public accounting of their incompatibilities with PyTorch (no one cares about Tensorflow anymore), even if the benchmarks show that their cards are worse. If you don’t measure in the public no one will believe your vague claims about how much you’re investing into the AI boom. Certainly I would like to buy an Intel Arc A770 with 16GB of VRAM for $350, but I won’t, because no one will tell me that it works with llama.