Of course, the most used stack is currently CUDA, which is proprietary and only works on Nvidia.
AMD made its own version called HIP, which translates quite directly to CUDA, but also runs on AMD cards via ROCm.
Intel has its own stack called oneAPI, which I think is based on SYCL, a Khronos standard (just like OpenCL/OpenGL/Vulkan/etc). I believe SYCL programs can also be ran on AMD and NVIDIA using third-party compilers/translation tools such as hipSYCL, and I think SYCL can also be compiled to OpenCL.
I recently also heard about efforts to support running HIP programs on SYCL, so hopefully soon GPU compute will be less vendor-bound, with CUDA translating relatively easily to HIP, and HIP and SYCL also not offering compatibility problems either way.
IMHO, AMD has made themselves irrelevant in GPU compute, and should work hard together with Intel to support SYCL on all GPUs and all popular OSs to enable developers and end users to use GPU for compute without being stuck with proprietary solutions.
That said, the build instructions for Blender say that you can't compile the GPU program on Windows. (I am not sure how this works exactly with JIT vs AOT compilation.)
AMD itself isn't working on SYCL to my knowledge, but you may be interested in hipSYCL which can compile SYCL code to HIP or CUDA. CUDA/HIP does have the advantage of a massive library of existing code and knowledge, and pre-existing libraries (e.g. cuPRIM, cuBLAS, etc) which you do not have on SYCL.
I also believe AMD has AOMP for OpenMP offload.
I do agree that AMD and Intel would be best off working together on GPU compute, so that they can say "look, our approach works on ALL vendors, not just us". However if Intel supports HIP on oneAPI, and maybe hipSYCL gets some sort of official support, we will be there already with both of their solutions.
As far as I can tell, HIP is still not generally available on Windows [0] or consumer gaming cards [1]. Not for developers, not for end users. That means hipSYCL can't be used by AMD users on Windows, and I don't see much value for Intel to support it either.
I don't think HIP is going to take any significant share of the market, why use AMD's platform when they so clearly have demonstrated that they don't really care about compute for more than a decade now?
I wish GPU compute could get to a maturity level where you could use any framework you wanted, compile it to SPIR-V, and then run it on any GPU on any OS. Let's say I write some image processing code, or ML inference, or audio processing, that I want to run on consumer's GPUs. What options are there today? Nvidia makes it easy, AMD makes it hard.
[0] https://github.com/ROCm-Developer-Tools/HIP/blob/develop/INS...
[1] https://github.com/ROCm-Developer-Tools/HIP/blob/develop/doc...
Also, my RX 580 is not officially supported, and doesn't work in Blender either (even when trying to compile it myself, as I got a compiler error related to LLVM limitations), but developing things myself seems to mostly just work so far. I also managed to port LeelaChessZero from the cuBLAS backend to hipBLAS without noteworthy problems (and gained a fair bit of speed over the OpenCL backend). (I still want to try and get it working with hipDNN, which I threw out at my first port attempt because hipify-clang got confused over various #IFDEFs but hipify-perl does not seem to have that issue).
Installing also hasn't been too bad for me on Arch Linux, I just installed an AUR package that pulls in most of AMD's binary packages (opencl-amd-dev) and now I have working hipcc and all that in /opt/rocm/.
I do agree that the situation could be far better, but if I can take a fairly simple CUDA app and port it in an afternoon to work on my consumer GPU, I don't think the situation is completely terrible either.
Also, as for various stacks, Close to Metal is very long ago. I've honestly never heard of Brook+ or Stream Computing. As for the others, I think AMD APP, HSA, ROCm, HIP and AOMP are all sort of part of the same stack. For example, my OpenCL driver which I believe is part of the ROCm stack identifies as AMD-APP. HIP and AOMP are then both like compiler frontends for the ROCm compute stack.
For the any-GPU any-OS SPIR-V wish, Vulkan or OpenCL are probably your best bet, although that won't meet "any framework".
Also, I don't buy that C99 is that much closer to the metal than C++17. From what I've seen, a lot of code still uses C-style anyway (e.g. C arrays). C++ just adds some nice stuff like namespaces. (I think a lot of the standard library would not work GPU-side anyway, at least with CUDA/HIP.)
Also, at least with CUDA and HIP, you need to manually malloc GPU memory, and explicitly launch GPU kernels (using <<<>>> syntax for CUDA or optionally hipLaunchKernel for HIP).
Ecosystem is the main problem of anything not CUDA or sort-of compatible like HIP. A ton production GPGPU code uses CUDA.
The only hope is that Intel pushes OpenCL as one of the selling points of their new Arc GPUs. They seem to be starting with the gaming market, however, which isn't very interested in computing. It could be that they have plans to attack the non-HPC market, in which case it would make sense to follow OpenCL.
At the same time, Intel is developing oneAPI, so it may make more sense to look at that instead of OpenCL.
From others' comments here, it seems there hasn't been much development since then in OpenCL land. I imagine details of device support might vary over generations of hardware and driver releases... I made use of half-precision float storage but ran compute at single-precision since my devices did not offer fast half-precision math.
This was small "hpc", i.e. a task that would run on one workstation or modest single socket or dual socket x86_64 server or VM. The most performance-sensitive aspect was that it was also used during the data loading/startup phase of an interactive tool. So a human user was impatiently waiting for results. The first prototype was just using Python numpy routines for convolutions, etc. I used OpenCL to get running time from many minutes down to tens of seconds and called it good enough.
I enjoyed that I could run the same code on a workstation GPU with the NVIDIA driver or via x86 multithreaded SIMD using the Intel driver. I did not do any real work with Intel nor AMD GPUs because of the hardware selection we had on hand. I also needed more than 6GB GPU RAM for it to be worth using. Even my Titan X with 12GB was only about 2-3x faster than using x86 SIMD for my problem, due to the complex tradeoffs between RAM and bus bandwidths for data transfers. This is after I'd already done some algorithmic optimizations to bring down the compute/IO ratio.
A big thing I hear others talk about is the rich development tools with CUDA and the relatively impoverished OpenCL tooling. I am old school enough to get along without it. I was able to compensate for limited Python OpenCL tools by doing some of my development and debugging cycle embedded in hacked up variants of my own viz tools. You might think of this a bit like debugging with printf, except my print statement could send a dense 3D array into an OpenGL-based renderer on my workstation.