A trip through the Graphics Pipeline (2011)
fgiesen.wordpress.com
fgiesen.wordpress.com
Updating the compute section to include WaveIntrinsics/subgroups, async compute queues, copy queue. Maybe even explore ML hardware (Tensor Cores)
This being said, you can learn WebGL and especially Three.js without knowing the pipeline as described there, but knowing it will help you understand some subtleties or if you ever want to move to a lower level.
DirectX12 and Vulkan have changed the pipeline by making it more flexible, but still largely follow the vertex -> geometry -> pixel pipeline.
Because the DirectX11 (and OpenGL) immediate-mode pipelines were used for so long, I expect that for the DECADES to come, the information in this blog will remain relevant (even if more efficient techniques advance the state of the art).
Assembly as the highest level abstraction. Anything equivalent to assembly anyway - a forth would be fine.
If i remember correctly he assembles a gfx card on a bread board. So its fairly fundamental
https://www.amazon.com/Designing-Video-Game-Hardware-Verilog...
That is abstractly what first principles graphics looked like some time ago - pixels were represented by a block of main memory, and the video system simply scanned that memory at the same frequency as the CRT monitor’s beam.
I’d say it’s pretty important to remember that what counts as “first principles” graphics has changed over time, and there was a time before frame buffers (the first Pong game was displayed on an oscilloscope) and the way we use framebuffers today is completely different - there is no more direct access to pixels in the same way old terminals used to work. The article here is detailing how we use a higher level API to draw higher level primitives, and the hardware takes care of the pixels.
I’d love to hear a little more about your goal... what’s driving you to study the old terminals?
Researchers like Blelloch have been writing in SIMD-parallel style "parallel Lisp" or "C-Star" for decades before CUDA. https://www.cs.cmu.edu/~guyb/papers/Ble90.pdf
EDIT: Woops, I mean CM2. CM5 was MIMD and a slightly different architecture. I've edited this post to say CM-2 instead, but my previous post about CM5 will remain in error.
Individual thread masking was available on CM-2, for all 4096 cores that were executing in SIMD-parallel.
CM2 was a 4096 x 1-bit SIMD processor. Very limited compared to modern GPUs, but the execution model that CM2 experimented with eventually became the modern GPU architecture. Yes, you can have effective execution masks even with only 1-bit cores.
EDIT: See the C-Star manuals, a "per-thread if statement" required a "where" construct instead of if. But yes, the "where" statement performed similarly to "CUDA if": http://people.csail.mit.edu/bradley/cm5docs/AReferenceDescri...
The MIMD thing that CM-5 did didn't seem to stick. CM-2's SIMD execution seems to be a bigger influence. But C-Star ran on both... and it was the high-level languages like C-Star or Parallel-Lisp that influenced GPU languages like CUDA.
> The where statement is involved with setting the context, a process known as contextualization. The context is a parallel boolean mask (i.e., each element of it is true or false) that controls the execution of parallel operations position by position. A different context is associated with each shape object, and the context associated with the current shape is always applied to operators.
A CUDA-block was roughly equivalent to a CStar-Shape. A CUDA-if is roughly equivalent to a CStar-where.
> Current GPUs also support certain amounts of fully independent thread scheduling, e.g., per-thread instruction counters, etc. Maybe that is beginning to leak out of the SIMT model.
NVidia was calling SIMT back in 2010, long before per-thread instruction counters were available in Pascal (GTX 10xx series of GPUs)
> Very limited compared to modern GPUs
Yeah, I'd expect the real true differences between CM5 SIMD and Pascal SIMT is just buried in the subtleties of those limitations compared to today's hardware, and that fundamentally they aren't super different ideas. In my mind, that doesn't mean that SIMT is meaningless or the same thing as SIMD.
> NVidia was calling SIMT back in 2010, long before per-thread instruction counters were available in Pascal (GTX 10xx series of GPUs)
Right; I'm saying the "SIMT" of 10 years ago has to do with masking and maybe other per-thread control (even though the Connection Machines might have had it as well), and that per-thread counters are now going beyond what we used to know as SIMT, perhaps deserving of some other acronym. Or maybe SIMT is becoming more appropriate and more differentiated from SIMD over time? Is it possible that CM5 should be called SIMT and we simply didn't have that acronym at the time?
Nope. I'm looking at C-Star through a historical lens, I was curious at the overall development of CUDA and was researching the "line of influence", so to speak.
> Right; I'm saying the "SIMT" of 10 years ago has to do with masking and maybe other per-thread control (even though the Connection Machines might have had it as well), and that per-thread counters are now going beyond what we used to know as SIMT, perhaps deserving of some other acronym. Or maybe SIMT is becoming more appropriate and more differentiated from SIMD over time? Is it possible that CM5 should be called SIMT and we simply didn't have that acronym at the time?
CM2 was just another SIMD machine, there were plenty of others before CM2 (its just that CM2 was one of the more popular ones). IIRC, Cray had similar SIMD machines that competed against them. The overall SIMD-methodology reaches back until the 70s at least, maybe earlier.
Just noting that CM5 was MIMD at the lowest level, trying to correct a mistake I made a few posts back.
------------
I think what "happened" was that Intel adopted SWAR: "SIMD With A Register" when Intel created the MMX / SSE / AVX instruction sets. (compared to dedicated SIMD-computers like the CM-2).
At some point, Intel's SWAR approach became colloquially known as SIMD (even though SWAR was much more limited compared to "proper" SIMD computers like the CM2). NVidia creates a new acronym called SIMT to differentiate Intel's SIMD (aka: SWAR) from a proper SIMD machine, even though the 80s-style SIMD was substantially similar to NVidia's new SIMT term.
the cm-1 and cm-2 both supported up to 64k single bit processors, and as many 'virtual' processors as you wanted until their memories got too small
the cm-5, while it did have a potentially mimd model, each mimd node was a sparc that was lashed to 4 simd vector units, with their own memory bandwidth. so you would be giving up a lot to just run on the sparcs. it was also pretty clumsy - I think there was just a message library one called from C?
I think that its not entirely fair to compare the CMs to the Cray vector products. although somewhat similar, the programming model for the Crays was basically loop mining fortran and the CM really presented a model of 2^n fine grained processors. the precedent I always heard was the Iliac
Note that Blelloch's favorite(?) primitive- the parallel prefix operation was used quite widely in the CM world, but equivalent 'horizontal' operations are pretty much completely lacking on the Intel/AMD instruction sets. Haven't looked at the GPU ones, but its absence is pretty frustrating
[0] https://excamera.com/sphinx/gameduino/
[1] https://excamera.com/sphinx/gameduino2/
It covers how graphics libraries like OpenGL go from triangles to screen coordinates, and how they "shade" pixels in those triangles to create an image.
For graphics programming in general Graphics Codex looks really good (I haven't done the course, but I know people that have). https://graphicscodex.com/