Optimization Techniques for GPU Programming [pdf]
dl.acm.org
dl.acm.org
As someone who has spent 80%+ of their time CUDA programming for the past 9 years (I wrote the original GPU PyTorch tensor library, the Faiss GPU library, and several things that Nvidia took and put into cuDNN), I found the most instructive, short yet "advanced" education on the subject to be Paulius Micikevicius' various slide decks on "Peformance Optimization"; e.g.:
https://on-demand.gputechconf.com/gtc/2013/presentations/S34...
(there are some other ones outstanding, think one was for the Volta architecture as well)
They're old but still very relevant to today's GPUs.
But I have no idea who employs those skills! Beyond scientific HPC groups or ML research teams anyway - I doubt they’d accept someone without a PhD.
My current gameplan is getting through “Professional CUDA C programming” and various computer architecture textbooks, and seeing if that’s enough.
CUDA is a polyglot programming model for NVidia GPU, with first party support for C, C++, Fortran, and anything else that can target PTX bytecode.
PTX allows for many other languages with toolchains to also target CUDA in some form, with .NET, Java, Haskell, Julia, Python having some kind of NVidia sponsored implementations.
https://developer.nvidia.com/language-solutions
While originally CUDA had its own hardware memory model, NVidia decided to make it follow C++11 memory semantics and went through a decade of hardware redesign to make it possible.
- CppCon 2017: Olivier Giroux "Designing (New) C++ Hardware”
https://www.youtube.com/watch?v=86seb-iZCnI
- The CUDA C++ Standard Library
https://www.youtube.com/watch?v=g78qaeBrPl8
It is also driving many of the use cases in parallel programming for C++
- Future of Standard and CUDA C++
https://www.youtube.com/watch?v=wtsnoUDFmWw
You will only find brief mentions of C here,
https://developer.nvidia.com/hpc-compilers
This is why OpenCL kind of lost the race, with it focused too much in its C dialect, only going polyglot when it was too late for the research community to care.
Unlike normal programming books, it talks a lot about how GPUs work and how the introduced techniques fit in that picture. It's interesting even if you are just curious how a (NVIDIA) GPU works at code-level. Strongly recommended.
So it's worth the update if you're interested in general NVIDIA GPU evolution.
Programming Massively Parallel Processors: https://www.youtube.com/watch?v=4pkbXmE4POc&list=PLRRuQYjFhp...
it's true - out of all of the "LEARN CUDA IN 24 HOURS" books, this is the best one. indeed this isn't one of those same books - this is a textbook - but at first glance it resembles them (at least the color scheme and the title led me astray when i first found it).
There's a book (https://metalbyexample.com/the-book/), but the author has put up a note that it's quite out of date. It seems the most up-to-date information is available in the WWDC videos (regarding e.g. Metal 3), but I'd really prefer something written. And Apple's documentation reads more like a reference material and is quite confusing when starting out.
https://www.amazon.com/Metal-Programming-Guide-Tutorial-Refe...
For the rest, yes, WWDC videos, samples, and then documentation, by this order.
Metal is actually one of the few new frameworks that happens to be written in Objective-C, with Swift bindings.
HN post here: https://news.ycombinator.com/item?id=37036058
In lieu of pointer chasing, hashing and the like, parallel operations on flat arrays are the way to maximize GPU utilization.
- compact a hash table (i.e., remove the empty slots)
- flatten a jagged 2D array
- rewrite a dense matrix in compressed-sparse-row (CSR) format
The outcome of a prefix sum exactly corresponds with the "row starts" part of the CSR sparse matrix notation. So they are also essential when creating sparse matrices.
AFAIK Thrust is intended to simplify GPU programming. It could well be that for specific use cases, in particular when it is possible to fuse multiple operations into single kernels, you could outperform Thrust.
Additionally Wgpu (the library) will insert fences between all passes that have a read-write dependency on a binding, even if there is technically no fence needed as 2 passes might not access the same indices.
Finally I know that there is an algorithm called decoupled look back that can speed up prefix sums, but it requires a forward-progress guarantee. All recent NVIDIA cards can run it but I don't think AMD can, so WebGPU can't in general. Raph Levien has a blog post on the subject https://raphlinus.github.io/gpu/2021/11/17/prefix-sum-portab...