> Is the new Xe architecture still SIMD?
SIMD is... pretty much all GPUs do. There's a few scalar bits here and there to speed up if-statements and the like, but the entire point of a GPU is to build a machine for SIMD.
NVidia started doing both around CUDA 3.0, whereas Khronos, AMD and Intel only started paying attention that not everyone wanted to do printf() style debugging with a C dialect until it was too late to get people's attention back.
From my understanding, the Khronos group realized OpenCL 2.x was much too complicated so vendors just weren’t implementing it, or only implementing parts of it, so they came up with OpenCL 3.0 which is slimmed-down and much more modular. It’s hard to say how much adoption it’ll get, but with Intel focused on DPC++ and oneAPI now, there will definitely be more numerical software coming out in the next few years that compiles down to and runs on OpenCL.
For example, Intel engineers are building a numpy clone on top of DPC++, so unlike regular numpy it’ll take advantage of multiple CPU cores: https://github.com/IntelPython/dpnp
Something like this also happened to OpenGL 4.3. It added a compute shader extension which was essentially all of OpenCL again, except different, so you had 2x the implementation work. This is about when some people stopped implementing OpenGL.
Khronos could have chosen only to add OpenCL integration, but OpenCL C is a very different language to GLSL, the memory model (among other things) is different, and so on. I don't see why video game developers should be forced to use OpenCL when they want to work with the outputs of OpenGL rendering passes, to produce inputs to OpenGL rendering passes, scheduled in OpenGL, to do things that don't fit neatly into vertex or fragment shaders?
By the way, the latest version is OpenGL 4.6, it is also available on the Switch.
DPC++ has more stuff than just SYSCL, some of it might find its way back to SYSCL standardization, some of it might remain Intel only.
OpenCL 3.0 is basically OpenCL 1.2 with a new name.
Meanwhile people are busy waiting for Vulkan compute to take off, got to love Khronos standards.
So far I am only aware of Adobe using it to port their shaders to Vulkan on Android.
I recently had a chance to learn the basics for a work project, never having touched the field before. I picked OpenCL, because I knew I was writing all my non-BLAS code myself, and there's no way in hell I'll voluntarily lock myself into a closed ecosystem. (PS: CLBlast, which is different from CLblas, is a joy!)
I was pleasantly surprised. I found OpenCL very nice to work with indeed! And my code runs on any modern GPU out there. I've tested it on Intel integrated GPUs, AMD GPUs, and Nvidia's fancy datacenter devices. And even CPUs. Seamlessly, through a runtime switch fully controlled by the application itself!
Now, could I have gotten more performance out of CUDA? Yeah, I estimate about a factor 2. For the cost of tying myself to a proprietary, locked in technology from a hostile vendor, throwing out two major classes of devices, and losing the ability to test out code anywhere. Not worth it.
I hope OpenCL has life in it still. The stuff I keep reading that CUDA is far easier to approach definitely did not ring true to this beginner.
Except NVidia Ampere (RTX 3xxx series) and AMD RDNA2 (Navi / 6xxx series) are both SIMD architectures with SIMD-instructions.
And the #3 company: Intel, also has SIMD instructions. I know that some GPUs out there are VLIW or other weird architectures, but... the "big 3" are SIMD-based.
> if you give one a fully scalar program it just needs to run a lot of copies of it at once.
Its emulated on a SIMD processor. That SIMD processor will suffer branch-divergence as you traverse through if-statements and while-statements, because its physically SIMD.
The compiler / programming model is scalar. But the assembly instructions are themselves vector. Yeah, NVidia now has per-SIMD core instruction pointers. But that doesn't mean that the hardware can physically execute different instructions: they're still all locked together with SIMD-style at the physical level.
You assume so, for Volta onwards that's not true.
Each SIMT thread on NVIDIA GPUs for Volta onwards has its own instruction pointer. And yes, it's scalar instructions at the ISA level.
https://arxiv.org/pdf/1804.06826.pdf
NVidia's SASS on Volta is pretty clearly SIMD-instructions. The PTX "virtual machine assembly" is fully documented. SASS is not documented, but it ain't a secret either. SASS follows closely with PTX, with exception of some "barrier bits" that seem to track dependencies (probably compiler-managed read/write dependency chains).
> Each SIMT thread on NVIDIA GPUs for Volta onwards has its own instruction pointer. And yes, it's scalar instructions at the ISA level.
When your "scalar" instruction executes 32-threads in parallel (subject to an execution mask), that's called SIMD.
The instruction pointer is to resolve deadlock conditions with locks, and the 32-wide SIMD cores will execute one-at-a-time to prevent deadlocks. But your goal as a GPU programmer is to get as much 32-wide execution going on as possible.
1-at-a-time serialization is very, very bad for performance. Its POSSIBLE to do on Volta, but highly recommended you stay away from that cornercase (you lose over 95% of your performance).
Each individual SIMT lane can actually diverge from the other on the underlying hardware. Active threads are dynamically mapped to SIMT units. You can use __syncwarp() to force reconvergence, to stall until all the threads in a warp are at the same location.
The disassembled code targeting the underlying ISA for a given CUDA kernel is also easily accessible through nvdisasm.
See figure 22 and figure 23 in: https://images.nvidia.com/content/volta-architecture/pdf/vol...
Thread divergence is bad: very very bad. Running 2x, 4x, or 32x slower (or less parallel technically). You want to avoid those situations as much as possible.
What NVidia noticed is that thread divergence is necessary for many classical locking algorithms: where serialized code must execute one-at-a-time to run an algorithm correctly. Under these conditions, tracking the instruction pointer, and turning off 97% of your cores to execute 1-at-a-time (instead of 32-at-a-time SIMD) is done.
That's WHY __syncwarp() exists, so that you can return to 32-at-a-time execution as soon as possible. Its not always possible for the compiler to figure it out, so the programmer can put a __syncwarp() as a compiler hint that 32-at-a-time is safe again.
(through "NVidia's SASS on Volta is pretty clearly SIMD-instructions")
Yes, thread divergence on GPUs can come with performance downsides and is a tool to use carefully.
Well, they execute 32-at-a-time SIMD, do they not?
There are tricks to split up the execution mask and execute 1-at-a-time, 2-at-a-time, 4, 8, 16, or 31 at a time based on if-statements or thread-locks or whatever. But a "SASS" instruction of "R1 = R2 + R3" is implicitly executed across a 32-wide warp if its execution mask is set as such.
----------
EDIT:
Lets put it this way: If you see a "SASS" assembly for R1 = R2 + R3 in isolation, how many adds take place on that clock tick?
Somewhere between 1 add, and 32-adds. No more than 32-at-a-time. Seems pretty SIMD to me.
I don't know why you continue to argue this...
if( threadIdx.x % 2 == 0){ // True for 16-threads
doA(); // 100 clock ticks
} else { // The other 16 threads
doB(); // 150 clock ticks
}
Lets say doA() takes 100 clock ticks, and doB() takes 150 clock ticks on NVidia Volta. How much time does the above code take to run?Answer: 250 clock ticks: doA() is run with 16-threads, then doB() is run on the 16-other threads afterwards.
That's how SIMD systems work. If this were non-SIMD 32-core turing machine, it'd take 150 clock ticks (And the threads doing doA() would have spend 50-clock ticks doing something else). The most important thing about SIMD from a performance perspective is that you "add" both halves of the branch from a performance point of view.
The underlying implementation is still just lane masking and walking down separate flow control branches sequentially of course.
The advantage of such a model however is that the complexity of this is abstracted away from even the compiler. This makes it a change in the programming model. That's why it's not _just_ named SIMD.
(I fear that this discussion went too far away arguing semantics... sigh)
That changes the history of the term SIMT. SIMT was first used for NVidia Tesla in 2006: over a decade older than the incremental changes made between Pascal -> Volta/Turing.
There has been no name change from Pascal -> Volta/Turing, as far as I'm aware. Furthermore, the PTX is substantially similar. The per-thread instruction pointers is pretty transparent in most code.
> The advantage of such a model however is that the complexity of this is abstracted away from even the compiler. This makes it a change in the programming model. That's why it's not _just_ named SIMD.
Have you looked at AMD's GPU ISA? Its just jump instructions, extremely similar to NVidia's PTX / SASS instruction set. AMD SIMD doesn't juggle execution masks explicitly either: its handled at a lower level (probably the decoder or something).
You still pull the execution mask for things like ballot instructions, but... both AMD and NVidia SIMD are pretty similar. (Similarly, NVidia PTX can still access the execution mask for ballot instructions as well)
------------
Although, I guess both of those GPUs can laugh at AVX512, where the execution masks are explicitly handled by the assembly programmer. But I don't know if explicit execution masks is necessarily a bad thing (its the job of the compiler instead of the decoder or whatever...)
> Unlike a vanilla SIMD machine, you can diverge more cheaply when you need to
If we take CM2 from 1985 as a "vanilla SIMD machine", it had execution masks and diverged extremely cheaply, just like modern machines.
C-Star and star-Lisp even had a programming model very similar to modern CUDA.
http://bitsavers.informatik.uni-stuttgart.de/pdf/thinkingMac...
Back then, you'd use "when" statements to do a parallel divergent branch, while "if" was only for uniform branches. But it wasn't like "when" statements were expensive, they just diverged.
> Furthermore, the PTX is substantially similar
PTX is just a (forwards-compatible) intermediate representation.
> The per-thread instruction pointers is pretty transparent in most code
Except control flow divergence, which is what changes there, yes.
> Although, I guess both of those GPUs can laugh at AVX512, where the execution masks are explicitly handled by the assembly programmer. But I don't know if explicit execution masks is necessarily a bad thing (its the job of the compiler instead of the decoder or whatever...)
That's a very good question. AVX-512 with its masking abilities was substantial progress over AVX2.
For a CPU, having secondary instruction flows just for the vector units just isn't a (reasonable) option though.
If there wasn't the 10nm issues, the next Xeon Phi would have been very interesting on that front. You might also want to look at the Fujitsu A64fx on the Arm side, used in Fugaku. (building a supercomputer with just CPUs, no GPUs)
We'll see what will be there in the future... will certainly be very interesting.
The way I would put this: the hardware is SIMD, and the number of operations executed per clock is the same as pure SIMD, but the independent instruction pointer per thread gives the scheduler considerably more flexibility, which is useful for avoiding deadlocks on synchronization, and also helpful for hiding memory access latency.
On top of that, there's an abstraction of a large number of scalar threads running on the hardware. It's a leaky abstraction, though, as performance really requires sympathy with the underlying SIMD reality, and also the subgroup operations also expose a good deal of the implementation details of that abstraction.
There was a talk by Andrew Lauritzen basically saying the same thing a few years back.
https://www.ea.com/seed/news/seed-siggraph2017-compute-for-g...
DX etc have been slowly exposing some of the underlying SIMD but it is still not really on the level of what is available in a full SIMD model like we have with AVX etc.
GPU languages like HLSL/GLSL are designed for ease of use and leave some performance on the ground.
However I would rather prefer some kind of "SQL" for GPU programming.