Gentle introduction to GPUs inner workings
vksegfault.github.io
vksegfault.github.io
Wrong, wrong, very wrong. No copy here. Mesa's RADV was developed completely independently from AMD, in fact it predates the public release of AMDVLK. It's also possibly the best Vulkan implementation out there. Valve invested heavily in the ACO compiler backend, so it compiles shaders both very well and very quickly.
To AMD's credit, they did help with documentation, questions, and much of the work was built on top of what was there for RadeonSI (OpenGL): shader compiler back-end, etc.
https://www.phoronix.com/scan.php?page=news_item&px=RADV-Rad... -> https://airlied.livejournal.com/81460.html
The status of some drivers is tracked here: https://mesamatrix.net/
There are state trackers for OpenGL, D3D9-10, OpenCL, and probably others. I guess one could make a Vulkan "state tracker", but it might not be efficient.
Gallium can be used on top of bare metal, but also CPU (LLVMpipe/OpenSWR), Vulkan trough zink, or now D3D12 thanks to Microsoft's recent work.
The description could still be improved, and I started commenting on how it could, but I realize this would warrant its own blog post.
It turns out these functions tied to how the GPU architecture runs the computing units in parallel. The computing units are arranged to run in parallel in a geometric grid according to the input data model. The computing units run the same program in LOCK STEP in parallel (well at least in lock step upon arriving at the dFdx/dFdy instructions). They also know their neighbors in the grid. When a computing unit encounters a dFdx instruction, it reaches across and grabs its neighbors' input value to their dFdx instructions. All the neighbors arrive at dFdx at the same time with their input value ready. With the neighbors' numbers and its own number as the mid point, it can compute the partial derivative using gradient slope.
Most instructions have no inter-dependency between cells. These instructions can be executed by a core as fast as it can on one computing unit until an inter-dependent instruction like dFdx is encountered that requires syncing. The computing unit is put in a wait state while the core moves on to another one. When all the computing units are sync'ed up at the dFdx instruction, they're then executed by the cores batch by batch.
If you want to execute more tasks than this, throw more cores at it.
Also the newer Volta architecture (section 3.2 in [1]) allows independent thread scheduling of threads across warps such that each thread can have its own program counter and stack.
There're certainly sync. See how barrier is used (9.7.12.1 in [1]) to synchronize threads.
[1] https://docs.nvidia.com/cuda/parallel-thread-execution/index...
Suffice to say, the Volta "Independent Thread Scheduling" is only used for compute shaders, not for vertex and pixel shaders.
You're right to note that different warps can run with different program counters, and when communicating across warps, those need synchronization. That's what those barriers are necessary for (see GroupMemoryBarrier and friends in the high-level languages). However, the definition of dFdX chosen by these languages basically requires that all threads participating in the derivative be scheduled within the same warp. dFdX will never synchronize or wait on another warp.
At the edges of a triangle, when it doesn't cover all four pixels in a quad the pixel shader is still evaluated on all four pixels to calculate derivatives and at the end the results for pixels outside the triangle are thrown away. Your pixel shader better not blow up outside the triangle or you'll get bad derivatives for the pixels that are inside the triangle.
As triangles get smaller, the percentage of pixel shader evaluations that are happening outside of triangles goes up. In the extreme case if the triangle covers only one pixel, the pixel shader is executed 4 times and 3 results are thrown out! This has started to become a problem for games with high geometric detail and it's one of the reasons why Epic started using software rasterization for their Nanite virtualized geometry system in Unreal 5.
You might say "why don't you just stop doing this approximate derivative thing so you don't have to execute pixels that are thrown away? Are derivatives that important?" The answer is that derivatives are used by the hardware during texture sampling to select the correct mip level, which is pretty much essential for performance and eliminating aliasing. So you'd have to replace that with manual calculation of derivatives by the application which would be a big change.
I find white papers quite good (although I admit there are many things I don't understand yet and constantly have to look up), but even these sometimes feel a bit general.
Hard to beat the actual technical references when you want something hardcore!
The actual CUDA documentation is good:
https://docs.nvidia.com/cuda/cuda-c-programming-guide/index....
In particular, "Section 5: Performance Guidelines": https://docs.nvidia.com/cuda/cuda-c-programming-guide/index.... gives a lot of micro-architectural details, including those nasty "bank conflicts" that people keep talking about. Honestly, the original documentation says it best with just a few paragraphs on that particular manner:
> To achieve high bandwidth, shared memory is divided into equally-sized memory modules, called banks, which can be accessed simultaneously. Any memory read or write request made of n addresses that fall in n distinct memory banks can therefore be serviced simultaneously, yielding an overall bandwidth that is n times as high as the bandwidth of a single module.
>
> However, if two addresses of a memory request fall in the same memory bank, there is a bank conflict and the access has to be serialized. The hardware splits a memory request with bank conflicts into as many separate conflict-free requests as necessary, decreasing throughput by a factor equal to the number of separate memory requests. If the number of separate memory requests is n, the initial memory request is said to cause n-way bank conflicts.
See? Its really not that hard or unapproachable. Just read the original docs, its all there.
AMD's documentation is scattered to the winds, but the same information is around. I'd say that your #1 performance guidelines are from the ancient optimization guide from 2015. Its a bit dated, but its fine: http://developer.amd.com/wordpress/media/2013/12/AMD_OpenCL_...
Chapter 1 and Chapter 2 are relevant to today's architectures (even RDNA, even though some details have changed). Chapter 2: GCN, applies for all AMD GPUs from the 7xxx series, through the Rx 2xx series, Rx 3xx, 4xx, 5xx, Vega, and CDNA (aka: MI100) architectures.
RDNA does not have as good of an architectural guide. Start with the OpenCL guide for optimization, and then "update" your knowledge with the rather short RDNA guide: https://gpuopen.com/performance/
For assembly language details:
* CUDA's PTX is basically assembly language: https://docs.nvidia.com/cuda/parallel-thread-execution/index...
* PTX is a portable assembly language: NVidia continuously updates their GPUs. Volta was very well studied by this paper: https://arxiv.org/abs/1804.06826. I'd suggest reading this AFTER you learn the basics of PTX.
* AMD publishes their assembly language for each generation. Vega is probably a good starting point: https://developer.amd.com/wp-content/resources/Vega_Shader_I...
Disclaimer: I have one un-merged PR in the gpu.js repo
[0]: https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Cent...
Oh, very much so. By way more than an order of magnitude. For a deeper read, have a look at the "architecture white papers" for Kepler, Pascal, Volta/Turing, and Ampere:
https://duckduckgo.com/?t=ffab&q=NVIDIA+architecture+white+p...
or check out the archive of NVIDIA's parallel4all blog ... hmm, that's weird, it seems like they've retired it. They used to have really good blog posts explaining what's new in each architecture.
You could also have a look here:
https://docs.nvidia.com/cuda/cuda-c-programming-guide/index....
for the table of various numeric sizes and limits which change with different architectures. But that's not a very useful resource in and of itself.
https://developer.nvidia.com/blog/cuda-pro-tip-nvprof-your-h...
For DL specifically, this article covers a couple of options that actually plug into the framework: https://developer.nvidia.com/blog/profiling-and-optimizing-d...
nvidia-smi is the core tool most folks use for quick "top"-like output, but there is also an htop equivalent: https://github.com/shunk031/nvhtop
A lot of other tools are build on top of the low-level NVML library (https://developer.nvidia.com/nvidia-management-library-nvml). There are also Python NVML bindings if you need to write your own monitoring tools.
They also i guess now have a web sever plugged into it which seems pretty cool
https://github.com/wookayin/gpustat https://github.com/wookayin/gpustat-web
I'm looking at graphics code, again, from a "I know enough C to shoot myself in the foot and want to draw a circle on the screen" perspective. It's hilarious how much "stack" there is in all the ways of doing that; I look at some of this shit and want to go back to Xlib for its simple grace.
GPUs solve problems of much larger scale, and so the stack has evolved over time to meet the needs of those applications. All this power and corresponding complexity has been introduced for a reason, I assure you.
There are interesting tasks in 2D graphics that may easily warrant GPU acceleration, such as smooth animation of composited surfaces (even something as simple as displaying a mouse pointer falls under this, and is accelerated in most systems) or rendering complex shapes including text.
A better explanation is that the main problems of processor design are that memory reads take 100-1000 times the time of an arithmetic operation and that hardware is faster when operation are run in parallel.
CPUs handle the those issues by having large memory caches, and lots of circuitry to execute instructions "out of order", i.e. run other instructions that don't depend on the memory read result or the result of other operations. This is great to run sequential code as fast as possible, but quite inefficient overall.
GPUs instead handle the memory problem by switching to another thread in hardware, and the parallelism problem by mainly using SIMD (with masking and scatter/gather memory accesses, so they look like multiple threads). This works well if you are doing mostly the same operations on thousands or millions of values, which is what graphics rendering and GPGPU is.
Then there are also DSPs that solve the memory access issue by only having very little on-CPU memory and either explicit DMA or memory reads giving a delayed result, and parallelism by being VLIW.
And finally the in-order/microcontroller CPUs that simply don't care about performance and do the cheapest and simplest thing.
What you're confusing is "Terascale" (aka: the 6xxx series from the 00s), which was VLIW _AND_ SIMD. Which was... a little bit too much going on and very difficult to optimize for. GCN made things way easier and more general purpose.
Terascale was theoretically more GFLOPs than the first GCN processors (after all, VLIW unlocks a good amount of performance), but actually utilizing all those VLIW units per clock tick was a nightmare. Its hard enough to write good SIMD code as it is.
There were two dimensions of "vector" -- whether you used XMM-style registers, requiring the compiler to auto-vectorize different operations. The old vector ISAs had things like dot product instructions. Those things are now all gone, because it turns out it's hard to optimize.
We in the industry call scalar-for-each-thread ISAs as "scalar ISAs" these days. For instance, Mali describes their transition from a vec4-for-each-thread-based ISA to a scalar-for-each-thread-based ISA as transitioning from "vector to scalar" [0].
[0] Compare 4:26 and 6:29 in the "Mali GPU Family" video on this page https://developer.arm.com/solutions/graphics-and-gaming/arm-...
Intel implemented SIMD using a technique they called SWAR: SIMD within a Register (aka: the 64-bit MMX registers), which eventually evolved into XMM (SSE), YMM (AVX), and ZMM (AVX512).
Today's GPUs are programmed using the old 1980s style / Connection Machine *Lisp, which is clearly the source of inspiration for HLSL / GLSL / OpenCL / CUDA / etc. etc.
Granted, today's GPUs are still SWAR (GCN's vGPR really is just a 64-wide x 32-bit register). We can see with languages like ispc (Intel's SPMD program compiler), that we can indeed implement a CUDA-like language on top of AVX / XMM registers.*
As far as I know, the first paper covering using SIMD for graphics was Pixar's "Channel Processor", or Chap [0], in 1984. This later became one of the core implementation details of their REYES algorithm [1]. By 1989, they had their own RenderMan Shading Language [2], an improved version of Chap, and you can see the similarities from just the snippet at the start of the code. This is where Microsoft took major inspiration from when designing HLSL, and which NVIDIA then started to extend with their own Cg compiler. 3dlabs then copy/pasted this for GLSL.
[0] http://www.cs.cmu.edu/afs/cs/academic/class/15869-f11/www/re... [1] https://graphics.pixar.com/library/Reyes/paper.pdf [2] https://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.21...
Wouldn't be graphics unless we overloaded a piece of terminology 20 times.
Also, how about Mesa3D copying Vulkan? It predates it.
It's always a bad sign when an author's exclamation has all the surprise factor of a tax code. As a previous poster said, "gentle" is relative.