Intel's next-gen Arrow Lake CPUs might come without hyperthreaded cores
tomshardware.com
tomshardware.com
It probably got a lot better with the i7, but experiences like that make me take new CPU features with a grain of salt. Hopefully someone can share more contemporary experience with hyper-threading.
With typical database servers, hyperthreading provides about 1.4x the performance "for free" because it allows execution units to do something useful while another thread is waiting for memory.
Where it isn't useful is typical game engines, except in some corner-cases as you've noticed.
Typical game engines like to have dedicated cores with the highest possible performance. Hyperthreading improves overall throughput at the expense of single-threaded throughput.
Similarly, game engines tend to run task-per-thread with different code, which tends to pollute the L1 code cache too much. So if two different threads are running on the same core but are doing different things, they'll fight over this precious resource.
Intel measured something like a +20% overall improvement averaged out across a wide range of workloads. A few had regressions, and a few (like ray tracing) had huge benefits.
IMHO, hyperthreading makes the most sense for servers, or laptops with a small number of cores (2-4). For high-end gaming, 16 cores with HT off will likely have the best performance.
https://www.hp.com/us-en/shop/tech-takes/why-should-i-upgrad...
Here's Reddit's take
https://www.reddit.com/r/explainlikeimfive/comments/66vy2r/e...
More recently we did some DirectX 12 MT-rendering tests; since there's more considerable overhead for breaking up the work into more batches so we wanted to be sure so we checked carefully.
On a i7-12700K we found that distributing Dx12 work to SMT or E-Cores had no benefit, even in a contrived rendering stress-test.
However, on an unnamed game console, we found that scheduling additional work to the SMT cores improved draw performance by maybe 7%.
YMMV but my impression is the workload needs stall on memory enough to get signfican gains by interleaving the ALU from another thread and this might benefit low-power systems like laptops or mobile phones more than workstations.
* hyperthreading is useless for purely-numerical code (the inner loop of math/science software), and for this kind of code it is also often harmful per the next point
* hyperthreading is harmful for code that is highly cache-sensitive within the scheduling quantum (though multiple cores may share caches too, and for the degenerate case of RAM-as-cache-for-swap people generally think in terms of "too many threads means more thrashing"; hyperthreading really isn't special there). This is more likely for "busy" tasks, but "background" tasks are likely to have no problem here.
* hyperthreading is beneficial for OO-like code that does a lot of pointer-chasing (which almost all software is, at least in part)
* hyperthreading might beat threads in separate cores for contended atomics
If you're trying to optimize particular code, it's generally better to eliminate the horrible memory patterns in the first place than expect a handwaved hardware "solution". Where it wins is for broader code that is too spread out for you to focus your attention on.
And of course, there is still a lot of code that is fundamentally single-threaded (at least in the critical path), and if allowed to hog the CPU will not see any benefit from hyperthreading.
For a long time hyperthreading was any easy win on desktop for all the various single-threaded background tasks. Now, with increasing core counts, specialized E- vs P-cores (and focus on power management in general), and Spectre/Meltdown, I'm not sure.
And when your very legitimate workload does stall, inevitably, it's nice if the CPU has another thread able to issue while it waits. In fact hyperthreading benefits almost all but the most heavily tuned workloads. Pretty much everyone stalls on a routine enough basis to make it worthwhile.
Now, in the post-spectre world where you need to spend transistors on firewalling/tagging to prevent information leaks, maybe it's not as big a benefit. But it's mischaracterizing history to claim that it was never worthwhile.
Just like HT.
CPUs target low latency (they switch often). GPUs target high troughput (they switch rarely, only when needed).
High troughput algorithms dont have problem with a lot of threads. Low latency algorithms have problem with a lot of threads (they need lot of cache memory because of constant switching).
https://www.hardwaretimes.com/intel-15th-gen-cpus-to-get-ren...
Getting rid of SMT gives you guarantees that certain resources are separated between software threads. That probably makes it easier for them to focus on making sure that other resources can be shared correctly.
You mostly want it for, say, low IPC server database workloads where the cores are largely just sitting there, twiddling their thumbs waiting for something from DRAM, where the very fancy branch predictors and cache hierarchies we have now can't help much.
Also, last time I checked (late penryn) the branch prediction and other elements necessary for good HT take up egregious amounts of die space -- which can probably be put to better use.
In the past Intel has claimed its not very much area, but (as you said) I am not so sure.
Look up some of the Jim Keller interviews on YouTube if you're interested in learning more.
edit: Instead running two hardware threads interleaved and sharing the resources on a single core, the implication is that you're turning it into a problem of scheduling work on separate cores with some provably separate resources. That probably makes it easier to have certain guarantees, and leaves you more time to worry about resources that you really do need to share but cannot afford to duplicate/distribute across the machine. Sometimes you really want two threads to share memory!
They beat on price and energy with no SMT.
This should be ok I assume in a similar vein if you can stuff as many cores as possible.
There is no such thing, basically core speed is a curve where they get more and more inefficient the faster they run, and the hard frequency limit is just some arbitrary point on that curve.
If a particular HT friendly workload gets more done at a specific clockspeed, that is a significant boost to efficiency even if it uses more power, as clockspeed power usage scales very non linearly towards the top.
That being said, I dont think you are wrong. In consumer devices, HT seems like a bad tradeoff for a bigger, more cache heavy core.
I think two cores that are 50% utilized burn through less power than one core that is 100% utilized but that also contains all the extra stuff it needs to pretend it's two cores. And that's the best case scenario. Quite possible that in the days HT was introduced, idle units weren't half as good at not consuming power as they are now.
The P-core tile is on the smallest, most expensive node, right? So Intel is probably getting more P-core tiles per wafer by removing SMT, and making up for it by giving you more of “other stuff” that comes from cheaper wafers with higher yield. As long as they are honest and you are getting performance that matches your expectations, this is fine.
This makes figuring out how many threads to run for optimal performance even more confusing.
If your threads are io-busy, pick thread counts based on io limits (depending on where on the throughput/latency spectrum you fall)
If you've got P and E cores, you might need to do some perhaps new stuff to balance threads --- the old way where an e core was about half the perf of a P core, so two threads on P and one on E was kind of even won't work without hyperthreads, but if you don't cpu pin your threads, maybe it works out. (Maybe the OS needs to do more work to track usage though)
Maybe the CPU can report how many threads it's capable of executing, regardless of how the cores are structured and their various capabilities, that'd be nice. I can imagine Intel might push this further, having 4 threads per core some day.
In the case where a core can run multiple threads, you have to benchmark to see if it makes more sense to run one thread per core (and possibly disable hyperthreading) or one OS thread per cpu thread. Sometimes there's a big difference, usually it's fairly small; and it can go either way. It's not unreasonable to run one thread per cpu-thread as a default, for cpu-busy code.
It's definitely not the case that if you have a core capable of running two threads that you lose half your throughput by only running a single thread. Depending on core design, you may or may not have a penalty by running single threaded and configured for two threads, if the core does a static partition of resources such as rename registers. But even in that case, losing half the rename registers is unlikely to reduce throughput by half.
https://www.notebookcheck.net/More-details-on-the-2024-Intel...
I am not up-to-date on Arc's architecture, but I am hoping that is in M Pro/256 bit memory bus territory. Big enough to run/finetune generative AI (especially with Intel's many own contributions to inference engines) with sufficient speed and plenty of RAM.
Really curious to see how the connectivity plays out. Raptor Lake has 16x PCIe 5.0 + 4x 4.0 on cpu, then 8x lanes to the chipset. But so far no USB4 on chipset.
Here it says Thunderbolt 4, which is 40Gb/s. Will those be through the DMI link too? They're also saying it has 80Gb/s, which probably is tied to the cpu. Now with USB/Thunderbolt you need more than some line mixing, so it seems more likely for the USB to be on cpu than chipset.
Then there is the whole security attacks that hyperthreading gives an helping hand.
another poster here commented on a 10% performance boost overall with HT on vs it off, that about matches up with my experience in an HPC-adjacent environment.