In fact, I have run tens of thousands of lines of essentially equivalent CUDA and OpenCL code (automatically generated) on the same hardware, and performance was in all cases very similar[0]. If anything, CUDA was actually slower than average (but in the cases I investigated, this was down to arbitrary differences like the CUDA compiler not unrolling some loops as aggressively and such).
[0]: https://futhark-lang.org/blog/2019-02-08-futhark-0.9.1-relea...
It's not. It uses LLVM, which can easily target AMD GPUs. (Whether the Julia folks have invested in making this work, I dunno, but it's not Extremely Hard.)
Understandably nvidia gives you the wrong impression.
The focus on CUDA comes from the fact that most HPC systems for scientific computing are using Nvidia GPUs. That is finally slowly changing.
Only when they started getting a beating of PTX bytecode and multi-language deployment on CUDA did they woke up and came up with SPIR (later SPIR-V) and SYCL, which still isn't widely deployed.
CUDA doesn't require rewriting code with ${OSS} framework of the year, every year. They need to earn that lock-in with future compatibility guarantees which none of OSS projects has.
> CUDA doesn't require rewriting code with ${OSS} framework of the year, every year.
How so? Change the GPU from Nvidia, and you are forced to rewrite code. That's the whole point of lock-in, it's a tax on developers. CUDA doens't guarantee you anything, if you don't stick with their GPUs.
Vulkan on the the other hand has conformance requirements.