NVidia started doing both around CUDA 3.0, whereas Khronos, AMD and Intel only started paying attention that not everyone wanted to do printf() style debugging with a C dialect until it was too late to get people's attention back.
From my understanding, the Khronos group realized OpenCL 2.x was much too complicated so vendors just weren’t implementing it, or only implementing parts of it, so they came up with OpenCL 3.0 which is slimmed-down and much more modular. It’s hard to say how much adoption it’ll get, but with Intel focused on DPC++ and oneAPI now, there will definitely be more numerical software coming out in the next few years that compiles down to and runs on OpenCL.
For example, Intel engineers are building a numpy clone on top of DPC++, so unlike regular numpy it’ll take advantage of multiple CPU cores: https://github.com/IntelPython/dpnp
Something like this also happened to OpenGL 4.3. It added a compute shader extension which was essentially all of OpenCL again, except different, so you had 2x the implementation work. This is about when some people stopped implementing OpenGL.
Khronos could have chosen only to add OpenCL integration, but OpenCL C is a very different language to GLSL, the memory model (among other things) is different, and so on. I don't see why video game developers should be forced to use OpenCL when they want to work with the outputs of OpenGL rendering passes, to produce inputs to OpenGL rendering passes, scheduled in OpenGL, to do things that don't fit neatly into vertex or fragment shaders?
By the way, the latest version is OpenGL 4.6, it is also available on the Switch.
DPC++ has more stuff than just SYSCL, some of it might find its way back to SYSCL standardization, some of it might remain Intel only.
OpenCL 3.0 is basically OpenCL 1.2 with a new name.
Meanwhile people are busy waiting for Vulkan compute to take off, got to love Khronos standards.
So far I am only aware of Adobe using it to port their shaders to Vulkan on Android.
I recently had a chance to learn the basics for a work project, never having touched the field before. I picked OpenCL, because I knew I was writing all my non-BLAS code myself, and there's no way in hell I'll voluntarily lock myself into a closed ecosystem. (PS: CLBlast, which is different from CLblas, is a joy!)
I was pleasantly surprised. I found OpenCL very nice to work with indeed! And my code runs on any modern GPU out there. I've tested it on Intel integrated GPUs, AMD GPUs, and Nvidia's fancy datacenter devices. And even CPUs. Seamlessly, through a runtime switch fully controlled by the application itself!
Now, could I have gotten more performance out of CUDA? Yeah, I estimate about a factor 2. For the cost of tying myself to a proprietary, locked in technology from a hostile vendor, throwing out two major classes of devices, and losing the ability to test out code anywhere. Not worth it.
I hope OpenCL has life in it still. The stuff I keep reading that CUDA is far easier to approach definitely did not ring true to this beginner.