Demystifying GPU compute architectures
thechipletter.substack.com
thechipletter.substack.com
GPU programming with CUDA and PTX feels like programming on a single core CPU without tasks and threads with deterministic behavior but in a multidimensional space. And every hour spent avoiding an 'if' pays off in terms of synchronization and therefore speed.
Along the lines of another comment on this post, part of the problem is the GPU compute model is a lot more abstract that what is presented for the CPU.
That abstraction is really helpful for being able to simply write parallel code. But it also hides the tremendous differences in performance possible...
I think the reason why Nvidia publishes these resources is because the GPUs are worth nothing if people can't get a reasonable fraction of the advertisable FLOPs with reasonable effort. CUDA wouldn't have taken off, if it were harder than it absolutely needs to be.
If I'm wrong, I'd be happy to learn more.
Not any contemporary mainstream GPU I am aware of. Sure, the way these GPUs are marketed does sound like they have superscalar execution, but if you dig a bit deeper this is either about interleaving execution of many programs (similar to SMT) or a SIMD-within-SIMD. Two examples:
1. Nvidia claims they have simultaneous execution of FP and INT operations. What this actually means is that they can schedule an FP and and INT operation simultaneously, but they have to come from different programs. What this actually actually means is that they only schedule one instruction per clock but it takes two clocks to actually issue, so it kind of looks like issuing two instructions per clock if you squint hard enough. The trick is that their ALUs are 16-wide, but they pretend that they are 32-wide. I hope this makes sense.
2. AMD claims they have superscalar execution, but what they really have is a packed instruction that can do two operations using a limited selection of arguments. Which is why RDNA3 performance improvements even on compute-dense code are much more modest. Since these packed instructions have limitations, the compiler is not always able to emit them.
Specifically, those structural diagrams of functional-units within SM's you see in the blog post? That comes from NVIDIA. And they explicitly state that they do _not_ guarantee that this is what the hardware is actually like. The hardware works "as-if" it were made up of this kind of units. And even more specifically - it's not clear whether there even is such a thing as a "tensor core", or whether it's just some hack for doing lower-precision FP math somewhat faster.
-----
Anyway, if the architectures weren't mostly _hidden_, we would make them far less mysterious within a rather short period of time.
They don't work like a von-Neumann machine, they act like one. Granted, we know a lot more about the inner workings of modern CPUs thn GPUs, but a lot of real-life work still assumes that the CPU works "as-if" it was a computer from the 70s, just really, really fast.
\rant I feel that generally people often have a notion of the "compiler is unbeatable at optimizing", which is completely false: You can always start from compiler output when optimizing and you often have easy options that are inaccessible for the compiler because of language semantics/calling conventions etc. (=> but hand optimizing at instruction level is time consuming, time is expensive, and what the compiler does is basically free ;).
Your software would not run next year if you directly targeted the instruction set.
NVIDIA does document their PTX instruction set (a level above what the hardware actually runs):
https://docs.nvidia.com/cuda/parallel-thread-execution/index...
Altough ofcourse CPU's instruction are also just a frontend api that behind the scenes is implemented using microcode, which probably is much less stable.
But the point is, if we could move one level 'closer' on gpus, just like we have it on cpus, it would stop the big buisness gate-keeping that exsists when it comes to current day GPU apis/libraries
That's not true for GPUs, the machine code changes very frequently. You feed your "binary" (PTX, ...) to the driver, and the driver compiles it to the actual machine code of your actual GPU.
The main difference is that with cpu, the translation unit is hidden inside the cpu itself. With gpus, the translation is moved from the device to the driver.
Old openGL code also will run in card that is newer then code itself.
The only difference is that with the cpus, it's open standard what is the instruction set, while on gpus, instruction sets are defined by third parties (DX12,Vulkan,OpenGL) while it falls to nvidia to implement them.
1. It's every two-to-three years.
2. It's not like they change into something completely different.
3. PTX is not a hardware instruction set, it's just an LLVM IR variant
4. I don't need to only target an instruction set directly, but it does help to know what instructions the hardware actually executes. And this is just like for CPUs (ok, not just like, because CPUs have u-ops, and I don't know that GPUs have those).
https://docs.nvidia.com/nsight-visual-studio-edition/3.2/Con...
https://docs.nvidia.com/cuda/cuda-binary-utilities/index.htm...
though there is no guarantee this is exhaustive, no opcodes either (though you could reverse engineer it using cuobjdump -sass and a hex editing like I've been doing). I'm pretty sure some of the instructions in the list are deprecated as well (95% percent sure that PMTRIG does nothing >Volta)
The first comment assures me "NVIDIA are really keen you understand their hardware, to the extent they will give you insanely detailed tutorials on things like avoiding shared memory bank conflicts."
I don't know who to believe!
Of course this doesn't apply to Nvidia, but MI-300 seems to be pretty viable for machine learning, so if you want things to change across the industry, more people need to put their money where their mouths are.
CDNA docs, for instance: https://www.amd.com/content/dam/amd/en/documents/instinct-te...
Go read NVIDIAs developer docs and AMDs training material for Frontier for some actually useful introductory material.
For a deeper dive, Jia et al.'s "Dissecting the NVIDIA Volta GPU Architecure via Microbenchmarking" is your goto for NVIDIA material, along with NVIDIAs own documentation. For AMD, you should probably read the CDNA 2 ISA docs
Sorry you don't think this is useful. However, as the post clearly points out it's not intended to make you a GPGPU programmer. That really isn't possible in a single blog post. Rather, it tries to give a general overview of how the GP part of GPUs works for the general reader, and to try to relate terminology used by AMD and Nvidia. I'm not aware of another source that tries to do this succinctly.
The end of the post gives lots of links for further reading for readers who want more.
And FWIW I have done quite a bit of GPGPU programming.
My main gripe is that the blog post does not seem to explain any of the terms it's inteoducing. The text keeps promising to explain concepts that aren't brought up again, and the main differences between AMD and NVIDIA are not discussed (shared memory, warpsize). I did appreciate the history in the introduction, that part was well-written.
If you want to explain the GP part of GPGPUs, I would suggest starting with the threadgrid and moving onto threadblock register footprints and occupancy, it's been a successful recipe for me when I've been teaching, and it keeps the mental load low.
(Does it end with:
"After the break we’ll discuss GPU instruction sets and registers, memory and software." (also side question, what break?) )?
Maybe people reacted this way because they assumed this was the full article
As to ML, it absolutely needs different hardware than your traditional GPUs. The reason we associate the two is because GPUs are parallel processors and thus naturally suit the task. But as we move forward things like in-memory processing will become much more important. High-performance, energy efficient matrix multiplication also requires very different hardware layouts than your standard wide GPU SIMDs.
Tomorrow's real-time rendering is all about RT, point clouds, microprimitives, and all that interesting stuff, the fixed function rasterizer simply won't be as useful anymore.
Ironically, iGPUs that use the system RAM bus or dGPUs that need two busses to get and put?
Is there a Northbridge and Southbridge anymore?
Kernels are executed at some blocksize, and each block is executed by a SM. The SM partitions each have a limited set of resources, and therefore the number of warps (or waves, AMD) that execute simultaneously should be tuned according to the resources available (registers, shared memory, etc.).
The keyword to search for here is occupancy, and related topics include register pressure/spilling, shared memory and L1 cache-size.
All extremely relevant.
And these aren't advanced concepts, it's just the fundamental programming model for GPUs. It's on the level of "Access memory in predictable patterns, ideally sequentially" for CPUs. Everyone knows now that CPUs like arrays and sequential access. And GPUs like interleaved acccess, ideally in sizes no larger than your shared memory.