Your software would not run next year if you directly targeted the instruction set.
NVIDIA does document their PTX instruction set (a level above what the hardware actually runs):
https://docs.nvidia.com/cuda/parallel-thread-execution/index...
Your software would not run next year if you directly targeted the instruction set.
NVIDIA does document their PTX instruction set (a level above what the hardware actually runs):
https://docs.nvidia.com/cuda/parallel-thread-execution/index...
1. It's every two-to-three years.
2. It's not like they change into something completely different.
3. PTX is not a hardware instruction set, it's just an LLVM IR variant
4. I don't need to only target an instruction set directly, but it does help to know what instructions the hardware actually executes. And this is just like for CPUs (ok, not just like, because CPUs have u-ops, and I don't know that GPUs have those).
https://docs.nvidia.com/nsight-visual-studio-edition/3.2/Con...
https://docs.nvidia.com/cuda/cuda-binary-utilities/index.htm...
though there is no guarantee this is exhaustive, no opcodes either (though you could reverse engineer it using cuobjdump -sass and a hex editing like I've been doing). I'm pretty sure some of the instructions in the list are deprecated as well (95% percent sure that PMTRIG does nothing >Volta)
Altough ofcourse CPU's instruction are also just a frontend api that behind the scenes is implemented using microcode, which probably is much less stable.
But the point is, if we could move one level 'closer' on gpus, just like we have it on cpus, it would stop the big buisness gate-keeping that exsists when it comes to current day GPU apis/libraries
That's not true for GPUs, the machine code changes very frequently. You feed your "binary" (PTX, ...) to the driver, and the driver compiles it to the actual machine code of your actual GPU.
The main difference is that with cpu, the translation unit is hidden inside the cpu itself. With gpus, the translation is moved from the device to the driver.
Old openGL code also will run in card that is newer then code itself.
The only difference is that with the cpus, it's open standard what is the instruction set, while on gpus, instruction sets are defined by third parties (DX12,Vulkan,OpenGL) while it falls to nvidia to implement them.