I meant dynamic (re)compilation yes - like adapting the stride in your kernel when you vary work group size for OpenCL kernels. I mean, I have something like that (a primitive self-tuning pre-run calibration step) and I only vary work group-, local- and vector size; and that's already a massive pain in the ass. It's ok on toy kernels but once you move beyond that, it just seeps into all your kernel code, making it almost into a meta-language. And then I'm not even talking about differences like using image types vs arrays for data that is 2d by nature. I don't see how I can generalize that; I just write various versions of my kernels. Which doesn't scale very much, to put it mildly.
My point is - was drawn to OpenCL for its 'portability' claim, and yes kernels will 'run' on various hardware, but with massive differences in speed. what good does that portability do me? My workloads (scientific computing, branch heavy) are different from the typical ML half- or single precision MulDiv()/linear algebra applications, so all the hand-tuned CUDA libraries aren't even my concern. The elephant in the room is that performance doesn't depend on this API or that; it's in how you tune your algorithm to the actual hardware you're running on. Which is the direct opposite of portability.
To come back to the post I was replying to - yes it's 'trivial' (I mean, a lot of work, but technically not hard, not to belittle your work) to compile almost any statically typed code into either SPIR-V or PTX or any other future format for that matter. But that doesn't mean that it will work 'well' (not even 'optimal') on other hardware. In the real world, you're almost always better off just spending a few hundred to get the same GPU as whoever wrote the kernels tested them on (or if you're running your kernels on existing clusters, to focus on optimizing them for what you know you'll be running them on).
Oh and all of that is just considering GPU's. I mean, when you read an OpenCL book they make it seem (in the first few chapters) that you don't even have to think about whether you'll be running on a CPU or a GPU. And then you accidentally use USE_HOST_PTR instead of COPY_HOST_PTR (or the other way around) in the wrong place, and all of a sudden your code is 10x slower than the sequential version of your algo even.
What I'm saying is - I no longer believe in easy to use abstractions for these purposes. If you want speed, you code to the metal, and/or you tweak your abstractions to your specific use case. Yes this is a lot of work. And if you don't need speed, you just throw in a few std::thread's here and there and call it a day.
But that's just my conclusion for my use cases.