This is explains the rise of ROCm and DPC++ as systems having equivalent higher level constructs to Cuda.
However almost no one cared, meaning AMD and Intel never shiped anything worth using, so OpenCL 3.0 is OpenCL 1.0 rebranded, and out of it, the C++ efforts were placed into SYSCL intead.
Which Intel picked up for their own DPC++ efforts, an additional tooling layer on top of SYSCL, meanwhile the only company selling usable SYCL developer experience is a former compiler vendor, that used to work with Sony in high performance compilers for the Playstation (Codeplay), as they pivoted away from console development.
Eventually Intel acquired Codeplay, and they are now the main supporters of One API tools, and the whole UXL efforts.
In the middle of all this, AMD decided to go with their own efforts.
It isn't only up to NVidia, when their competition can't get their story straight for decades.
If anything the situation is that Nvidia is far-sighted, developing hardware and software for general purpose GPU computing and the other chip manufacturers still think like traditional (non-CPU) chip manufacturers - they just want to sell a best-chip for a single purpose. This explain why they throw random things against the wall to see if they stick rather than choosing a general purpose and committing to it indefinitely.
People have even started using Vulkan in its place for some GPGPU, because at least Vulkan has good drivers everywhere, being a low-level API.
The problem is that the drivers are merely OK. Presumably if you're using OpenCL you care about the performance (otherwise why would you??) and since that's the case, it's the best on no platforms, and there are alternatives for any set of platforms that do better.
I think OpenCL is sadly on its way out, and it's mostly Apple's fault (and Nvidia a little). Vulkan compute is much more interesting if you're looking to leverage iGPUs/mobile/other random CPUs.
If you're targetting workstations/server workloads only, it makes sense to restrict yourself to a subset of accelerator types and code for that (eg. Torch or JAX for GPUs, use highway for SIMD, etc.)
I'm definitely not saying OpenCL is any sort of a reasonable default for cross platform GPGPU work. In truth, I don't think there is any reasonable "general" default for that sort of thing. Vulkan has its own issues (only works via a compatibility layer on MacOS, implementation quality varies widely, extension hell, boilerplate hell, some low level things are just impossible, etc.) and everything else is a higher level approach that can't work for everything by definition.
It's a pretty sad situation overall and every solution has severe tradeoffs. Personally, I just write CUDA when I can get away with it and try to stick to OpenCL otherwise, but everyone needs to make that choice for their own set of tradeoffs.
But the driver implementations inconsistency, version support issues, etc. meant people used CUDA instead.
I agree Vulkan has its own issues, and having written some MoltenVK stuff, you clearly know the quality-of-life pains in developping with it. That said, at least from the user side it works and performs well.