While most programs are not embarassingly parallel in nature, I suspect the latter plays a bigger role than most of us tend to admit.
Right now we are two steps removed. There aren't even many good data structure implementations out there. There is moodycamel::concurrent_queue and different memory allocators like jemalloc and the new rpmalloc, but the vast majority of programmers likely don't even realize that malloc locks and kills concurrency. Maybe they grab something from boost that just surrounds a data structure with a mutex, which is lazy bullshit.
I actually think the majority of software that runs slow can be made to be sped up nearly linearly using dozens of threads, but the few niche libraries that help are overly complicated to compile and use, bloated with enormous dependencies and end up being a minefield of usage. At the moment many people's idea of multi-threaded is mostly fork-join techniques like openMP and that will never scale to using all cores throughout the whole program, since is one very narrow technique.
If you think about the heart of the problem, it is really synchronization, and that only needs to happen in certain places in programs. Anything you can break up into pieces that can be transformed independently can be leveraged for concurrency, which also implies that independent stages in a pipeline can be concurrent as well.
Otherwise I think the rest of Erlang's performance is because it's optimized for reliability and distribution rather than raw performance.
One is promises (async/await). This takes the burden off the programmer for explicitly managing concurrency, and potentially allows lazy evaluation in many cases (if non-side-effecting). I've never worked in Lisps but it seems like a viable model in many ways.
The one I wanted to bring up was probabilistic programming. Imagine a tree where there's not just one place to put something, there's many, possibly distributed. As such you greatly reduce the chances of contention from multiple threads. They would also probably (almost necessarily) reduce the amount of rebalancing work.
GPUs seem like the perfect model here in multiple ways, not only do you have the massive degree of threading that surfaces any problems in such things but you also have the degree of threading to deterministically search these kind of structures. GPUs also highly discourage things like CAS or mutexing that we need to move away from.
(PS nice username)
GPUs are useful for the kinds of parallelism that has always been known to be easy - fork join with all threads writing to new memory. There is no mystery to how to do this.
> CAS or mutexing that we need to move away from
compare and swap is the backbone of lock free programming. Mutexes to me seem to rarely be ideal, but there is no getting away from compare and swap as far as I know.
The nice thing is, I can throw a new GTX and a PCIe SSD in and that'll have quite a bit of benefit. Custom PCs are awesome!
The single thread performance on the 6950X is pretty low, and in fact pretty much on par with that 4771. If you want higher single-thread performance you have to opt for fewer cores at a higher clock rate, something like the i7 7700K, which would definitely perform significantly better than the 4771. The 6950X really only makes sense if you have a workload that needs those 10 cores.
See https://www.cpubenchmark.net/singleThread.html for the comparison.
Broadwell is kind of an exception here since it clocks so poorly both in the small-chip and HEDT form. But for example, 4.7 GHz is a good OC for Small Haswell and the Haswell-E HEDT chips will usually go to at least 4.5 GHz.
(of course the major caveat here is power - a highly OC'd 8-10 core chip can easily pull 200-250 watts package TDP on a normal workload, 300+ watts as measured at the wall. For an AVX-heavy workload like Prime95, add another 100 watts package TDP.)
So overall the penalty for using HEDT chips is nowhere near as big as you are implying here. It's something on the order of 4% loss of single-thread performance to go from 4-core to 8-10 cores. It just doesn't come that way out of the box - largely because anyone who buys a 10-core gaming chip for $1700 knows exactly how to do all this.
The bigger problem is that the HEDT chips are 2 generations behind the smaller chips. Broadwell vs Broadwell-E, the numbers look fine, but Kabylake has some IPC improvements and clocks significantly higher than Broadwell does.
This is one of those things that actually can be blamed on "Intel being lazy because they have no real competition". They could certainly push the HEDT chips out a little faster instead of waiting to work out all the bugs in the smaller chips, it would just be a bit riskier for them.
Think about that: a 4 back your PC started in 10 seconds and 2 years back it's starting in 5 seconds, that's 5 second difference that you feel.
Currently the PC's start at 2.5 seconds and the speed difference is much less, so it feels like the power is not progressing that fast.
I have recently "upgraded" my desktop with an SSD (but have an old CPU i7 4780k or something) and it loads in a few seconds. But I don't want to upgrade my CPU just yet as it will only make everything only 1 second faster, not worth 300 USD I think.
But if you do a lot of data processing, software builds, ... You can increase the speed much more, and then it should matter.
Then you have improved architectures, memory/cache access, inter-core communications... there are more ways to improve a chip.
I agree with your snetiment, for general workloads. I have a 4 year old Haswell and it's working like I imagine the i5s today would perform on a basket of workloads.
Thanks!
Equating the single core speed of Intel's latest CPUs and moore's law is just another case of simple, easy and wrong.
> For a while they were able to use higher transistor counts to get more instructions per clock, but even that strategy is running out of steam
On a single core, not over a whole processor, as can be seen by the fact that there are 22 and even 72 core processors available now.
> The big news about Ryzen is that it was finally able to catch up to Intel in that regard.
This has nothing to do with Moore's law, it is a result of processor architecture.
> On a single core, not over a whole processor, as can be seen by the fact that there are 22 and even 72 core processors available now.
A worthwhile development, but not a panacea. I'd suggest you become familiar with Amdahl's law: https://en.wikipedia.org/wiki/Amdahl%27s_law. There's also some give-and-take between core count and individual core performance.
> This has nothing to do with Moore's law, it is a result of processor architecture.
But Moore's law is an enabler of more sophisticated processor architecture.
Once again, there are still transistor density improvements.
> A worthwhile development, but not a panacea. I'd suggest you become familiar with Amdahl's law
I'm not sure what you are trying to say here. Are you going back to saying computers aren't getting faster and then trying to say that moore's law is therefore dead when we've already established that it was never about speed?
Also what is your point about Amdahl's law? I work with lock free concurrency all day every day and if you are trying to say that more cores doesn't mean more speed because of Amdahl's law, that is pretty shaky ground to say the least. For some reason people seem to assume that if the software they use isn't using multiple threads effectively that it must be impossible. This is FAR from the case. Most programmers don't know how to work with concurrency and most software is trapped in legacy architectures that make parallel computations become painful changes.
> But Moore's law is an enabler of more sophisticated processor architecture.
Again, this tangential and not even really true. AMD's CPU architectures suffered while their GPUs performance and manufacturing process has remained competitive.