But then again, number of cores is really a crude measure. We need to measure what those machines can really do...
P.S. Surprised I am being downvoted. If you only want to compare number of cores then I was implying that 80 is a better number than 40 to compare with 128. It gives you a rough idea of parallelism even though obviously HT does not give you 2x available cores all the time.
I also added a caveat that that number of cores is a crude measure anyways.
If you're chasing dependent loads around memory, like traversing a long fragmented linked list due to garbage code, you will get an absolutely perfect 100% speed up.
If you're already saturating the ALU on one thread, adding another to the core will probably slow you down with context switching and cache contention, and indeed well written numeric simulations on supercomputers often turn it off entirely.
However, most software we run resembles the garbage pointer chasing variety more than the finely tuned numerical variety.
To be fair, crappy pointer-heavy data structures is not the only possible reason. Some useful algorithms are inherently serial.
For instance, all streaming parsers, compressors, or cryptography are inherently serial algorithms. An implementation gonna be relatively slow not because of RAM latency, but due to continuous chain of data dependencies between sequential instructions.
It’s technically possible to implement single-threaded code to handle multiple streams concurrently. Practically, more often that not it’s prohibitively complicated to achieve. However, doing that in hardware with two hardware threads running on the same core is way more manageable in terms of software complexity.
I actually experimented with parallel fetching of multiple values out of a large persistent vector years ago, and saw nearly linear speed up for up to four fetches in parallel.
The code was awful and it would need to be a compiler generated thing for sure.
C++ templates technically work, Eigen [1] is made entirely of them and the performance is pretty good. However, writing and debugging code in that style is hard, and the code is unreadable.
I tried T4 [2] to generate AVX intrinsics-heavy C++ code. At least the programming language is sane as opposed to C++ template metaprogramming. However, the code generation adds friction. Also, IDE integration is less than ideal, apparently Microsoft only cared about consuming T4 from C# projects.
I basically gave up, and usually writing normal C-like code with SIMD intrinsics, with just a few preprocessor macros, and no template metaprogramming. Unrolling things manually when necessary. At least the code is readable this way, and not too hard to debug. The compilers have these non-standard keywords/attributes to forcefully inline functions, they allow to reuse code instead of copy-pasting.
[1] https://eigen.tuxfamily.org/index.php?title=Main_Page
[2] https://en.wikipedia.org/wiki/Text_Template_Transformation_T...
Actually, for my specific use case using AVX512 gather intrinsics would probably have worked OK and been cleaner - although I've heard they're pretty slow.
See https://www.intel.com/content/www/us/en/developer/topic-tech...
So the level at which your Intel CPU is "crippled" would depend on how old it is, I guess.
A lot of expert level folks on here were saying that the architecture surrounding speculative execution would need a total removal or reengineering to fix it.
If your using the CPU core as efficiently as possible, you'd see no benefit from hyperthreading.
If your using it very poorly, you'd see a massive 2~x benefit.
That said I did switch from Intel over to an Apple M1 anyway.
What I would find really, really interesting: a "single-process" compiler that has a global in-RAM cache for all source contents and intermediate outputs and can avoid the overhead of child processes... basically a model like Webpack or Parcel that has an inotify watcher and is constantly running. The JS world had no other choice with NodeJS/npm all but forcing the tooling to adapt to a lot of incredibly small source files, it's time for the "classic" world to adapt.
I find this statement tough to agree with. It depends on the kind of work your processors are doing: it is IO heavy work or are you running Math computations? There are other axes basically related to how much work can be done by the current thread while the other thread stalls waiting for data to be fetched (or other non-parallelizable dependencies to be available).
So if you're getting a high benefit you shouldn't necessarily feel embarrassed!