Nuts and Bolts of Multithreaded Programming (2011)
software.intel.com
software.intel.com
I think it gives a great high level overview of how parallel programming works. Great starting point for those interested in learning more about HPC systems!
I think the focus was on vectorisation rather than task parallelism, though. And it's been abandoned for years.
I'm not sure Intel's sustainability is uniquely dependent on parallel programming. In fact, given that Intel processors have the best sequential performance, one could even argue that it's to their advantage that parallel programming remains awkward.
http://chimera.labs.oreilly.com/books/1230000000929/index.ht...
> As Intel puts two cores on a single piece of silicon and as multi-socket systems continue to grow, shared memory systems are going to become the norm.
Clearly this was originally published quite a bit earlier.
And surprisingly, there is no mention of GPU.
Almost as if they want you to think parallelism can only be achieved through CPUs (cores, and threads) but don't want to admit it in the title.
Also note that it was written in 2011, not sure if GPU based HPC was as common then but I could be wrong.
It did acknowledge the title, and contrasted it with the tone of the article (the first half at least) which is way broader minus a mention of GPUs (the basics of parallel algorithms, parallel APIs).
SIMD = Single Instruction Multiple Data, meaning the same instruction being applied to multiple different values simultaneously. That's exactly what GPUs do.
That being said, I understand you only wanted to point out the error in the upper post.