Is Parallel Programming Hard, And, If So, What Can You Do About It?
kernel.org
kernel.org
Massive-scale parallelism uses different data structures, algorithms, and programming models not discussed.
There's a handful of papers on attacking the 13 (formerly 8) dwarf problems http://tom.scogland.com/pubs/pdf/feng-ocd-icpe12.pdf
I don't have any specifics on Kaveri memory model but there's lots of small reviews: http://www.tomshardware.com/reviews/a10-7850k-a8-7600-kaveri...
compiler research and thread debugging: http://iacoma.cs.uiuc.edu/iacoma-papers/Illinois_parallelism...
Introduction to High Performance Scientific Computing - Victor Eijkhout with Edmond Chow, Robert van de Geijn
The memory hierarchy is our enemy here: the reason GPUs have done so well is that they schedule memory just as much (if not more) as they do computation. If you are going to go through the trouble of coordinating threads to share caches (and this is possible at all), you might have a GPU-friendly problem.
Unless, maybe, you mean the communication costs are the real bottleneck in such cases? In which case I don't see the relevance of the GPU angle.
DNN training is one problem where the GPU solution vastly outperforms the distributed HPC solution.
2. GPUs have less "close" memory (registers/shared memory/cache) per vector lane than current CPUs. This means that to get high efficiency, you have to find fine-grained parallelism without additional overhead. This is hard and it is common to see GPU algorithms make more round-trips to global memory than analogous CPU algorithms, eating into any benefits in raw bandwidth.
3. GPUs have a relatively narrow range of problem sizes in which they perform well. You typically have to use a significant fraction of device memory to expose enough parallelism to keep all the cores busy, and yet the device has limited memory compared to the CPU. On a GPU-heavy configuration, you have placed 90% of your compute next to 10% of your memory, connected by a straw (PCI bus) to the rest of the memory and the network. That is not a recipe for a versatile machine. CPUs give you vastly more flexibility in turn-around time (e.g., strong scale at >50% efficiency over a factor of 1000 as compared to 10). GPU performance results usually choose a problem size that fills device memory, but science/engineering is often not that convenient.
4. Even for problems in which GPUs perform optimally (like DGEMM or HPL), the ratio in energy efficiency is only 2x. See http://green500.org for example. Note that Blue Gene/Q is a CPU architecture that delivers the same energy efficiency as the GPU-heavy Titan. Also note that Haswell improves Intel efficiency by 2x over Sandy Bridge. The 1000x myth needs to die.
5. Enterprise GPUs (those with ECC) and Xeon Phi (MIC) are expensive ($3k-4k MSRP) relative to CPUs, and still need a host in almost all configurations. In performance tests, normalize-by-shrinkwrap needs to die. Normalize by total acquisition cost or by total energy consumption (always include the host, memory, network as applicable).
2. Every GPU generation adds more things like cache. The story changes every two years in favor of GPUs.
3. My colleague trains huge multi-gigabyte models on GPUs, so its not impossible. Terabytes of data is still the domain of MPI and increasingly MapReduce.
4. This is really not a concern for us, nor anyone who is in it for the performance (6 hours vs. 6 days).
5. The Phi is still kind of a joke. Tesla is quite competitive in terms of pricing.
Modern individual machines give us all kinds of opportunity for parallel execution—even my phone has vector instructions and two CPU cores. Modern Intel CPUs have 256-bit wide vectors, meaning you can do 8 floating point operations in a single instruction, assuming you can phrase your problem in the right way. If that's not enough, you can throw a pile of these cores at the problem, assuming you can coordinate work between these threads.
My experience has been most programmers struggle hard to build multi-threaded applications, countless hours lost to tracking down race conditions, deadlocks, and unexpectedly bad performance, which is a shame.
Scanning through this book, it doesn't appear to cover vector instructions, but it does look like it covers coordinating multiple threads in amazing depth. I feel like I already have a decent handle on multi-threaded programming, but I will be reading this book, and I'll probably learn plenty.
A lot of it is caused by tools which are by design prone to such conditions and unless you consciously and meticulously follow very strict guidelines, such problems are inevitable. But some tools try to solve it with different design approach (such as Rust language for example) which prevents many of those potential pitfalls implicitly. I wonder if the book is focused on shared memory and locking only, or covers broader range of methodologies (at the first glance it's mostly about classic shared memory and mutual exclusion approach).
Does it mean, you need greater knowledge? (communication models, deadlock/livelock, etc)
Does it mean, you need greater attention to detail? (mistakes that wouldn't matter become serious)
Does it mean, you need greater working memory? (remembering what needs locks, and what already has locks; in addition to more ordinary side effects and possible exceptions etc)
Does it mean, you need greater fluid intelligence? (reasoning about which locks can be safely composed, or which memory transactions will have too much contention)
Each release is typeset both single-column and double-column. Single column works well for the larger-format ebook readers, and double-column works well for laptop/desktop use and for hardcopy.
A number of people have reported good results with the single-column format on higher-end ebook readers. The single colume version of the first electronic edition may be found here: http://kernel.org/pub/linux/kernel/people/paulmck/perfbook/p...
And for the final nail in the coffin you're probably asking your program to do something that violates the known laws of the universe in regard to the speed at which information can travel.