The problem with parallelisation
cartesianproduct.wordpress.com
cartesianproduct.wordpress.com
But on a GPU, it is typical to have well over 1000 cores using a different programming model that accounts for memory access and grouping threads together to accomplish shared work and shared memory accesses.
So, the same work is now being done with slightly different algorithms that exploit this programming model.
I rewrote a toeplitz matrix solver not so long ago; it was pretty fun! Treat conditional statements as plagues, and don't worry as much about early stopping... Every example takes the worst case time to handle, but you can do thousands of them at the same time. Ended up getting about a 1000x speed up against the baseline.
If your algorithm requires any serial coordination between execution units, you're hosed, whether you're using CPU, GPU or really anything that can be used for computation.
There are a set of algorithms that are "embarrassingly parallel", and that's where for example many graphics tasks fall. Same is not true in general.
Is it something software can do anything about?
ex: df.groupby('x').agg(my_weird_symbolic_func) can do well in rapids.ai and badly on AVX.
Note the first line of the text specifically references "general computing."
If you're looking for advancement in computing power in the same programming paradigms and hardware styles you've always used, you will eventually be disappointed.
We've had GPUs for ages. Programming for GPUs is common, though I would suspect has significant innovation remaining.
As a comparison between GPUs and CPUs, each "CUDA core" is approximately equal to a SIMD lane on an x86 core. So 56 cores, times 16 lanes (AVX-512) per core, is 896 cores.
It's not really, that's why Larrabee failed.
It has bothered me a long time how the whole industry has fallen to GPU manufacturer marketing speak.
That Nvidia Titan RTX thingie is a monster, but it has only 72 SMs (Streaming Multiprocessor), the closest equivalent to a CPU core. That is, SM is the smallest unit that can individually execute a true branch. Just like a CPU core.
Not counting hyperthreads here, GPUs have a metric ton of them, because they have to do useful work while waiting for very high latency memory. CPUs need much less of them to cover "bubbles in the pipeline".
"Like Volta, the Turing SM is partitioned into 4 sub-cores (or processing blocks) with each sub-core having a single warp scheduler and dispatch unit"
"Like we saw in Volta, these changes go hand-in-hand with the new scheduling/execution model with independent thread scheduling that Turing also has, though differences were not disclosed at this time. Rather than per-warp like Pascal, Volta and Turing have per-thread scheduling resources, with a program counter and stack per-thread to track thread state, as well as a convergence optimizer to intelligently group active same-warp threads together into SIMT units. So all threads are equally concurrent, regardless of warp, and can yield and reconverge."
GPU 'cores' are still not as general as a full CPU core but they are a fair bit more flexible and sophisticated than a SIMD lane.
https://www.anandtech.com/show/13282/nvidia-turing-architect...
So what are the sub-core limitations execution flow wise? Or is it truly generic?
Can there be execution level conflicts? Like execution flow could be taken, but execution resources constraining (stalling) it somehow?
I'm only talking about "per clock" view, not about hardware threading. Sometimes these terms are mixed in a bit hard to decipher fashion.
I suspect some of these changes were to make raytracing more efficient (and the Anandtech article suggests as much too). A common problem trying to implement a GPU raytracer on older architectures is performance tanking when rays diverge too much. NVIDIA's Optix did some complicated magic to try and regroup divergent threads but having direct hardware support for it probably makes the raytracing GPU code both simpler and more efficient.
I thought they used a simpler barrel-processor model to handle that. SMT seems a bit more sophisticated.
AVX-512 has 32 SIMD registers.
Where does that 250x slower come from? That doesn't seem a very reasonable assumption.
So, basically, any particular problem will only go n-times faster thanks to Ahmdahl's law, but with the advent of more CPUs, we feel less restricted in our problems, and as a result work on bigger data sets.
After all, there already exists plenty of software (including some I wrote) which is limited by memory bandwidth, or by latency.
Also, all of those 1000 cores would compete for the same main-memory bandwidth, so are much more likely to be bandwidth starved.
SIMD instructions and GPGPU has come a long way and while clock rates have somewhat stagnated, we can solve many problems on consumer hardware today that were impossible in Pentium 3 days (big ML models, augmented reality, live ray tracing, ...). It was a painful transition from existing computation models, but it has been a gigantic success imho.
[1] https://battlepenguin.com/tech/microservices-and-biological-...
[2] https://spectrum.ieee.org/semiconductors/processors/cerebras...
We tend to think that replacing a O(n) algorithm with an O(n ln(n)) algorithm is a bad idea. But if it lets you spread out your problem over multiple threads it might very well be worth it. Within a certain range the power and silicon required to execute sequentially at a certain speed is, to simplify hugely, a bit more than the square of that speed so there's a lot of potential for improvement.
Certainly anything a human brain can do can be done in an embarrassingly parallel way which ought to be of some comfort.
This essay concerns 1024-core chips. ("When I started the PhD as a part-time student in 2012 the firm expectation in industry was that we would by now be well into the era of 1024-core chips. That simply hasn’t happened because, at least in part, there is no firm commercial reason for it")