Intel Ditching Hyper-Threading with New Core I7-9700k Coffee Lake Processor
wccftech.com
wccftech.com
(FWIW, this comes from people who tend to know what they are talking about. But also people who value security over everything else)
At least there's no proof so far
I just wonder if I'm in the minority running more than one app at a time. Specifically, a dozen apps that are never closed and just stay in the background until I need them. Even more specifically, listening to music on YouTube while working in the IDE, wherein the browser pops up regularly on the CPU usage chart.
For the same reason, the norm of 4 or 8 gb RAM is baffling to me.
In general if you only care about raw performance you'll run one heavy app per machine. You won't use the same server to render video and build your code at the same time, it will be more efficient to use two different servers for that.
Most of the time, threads and processes are in the blocked state. For example, waiting for mouse movements, or network traffic.
Check your CPU utilization, I bet you its below 20% if you have anything close to a modern processor. Even with 30+ tabs open
I do care about performance at work because I'm using it 8 hours a day but I wouldn't min an i5 at Home at all. Don't think that any normal person needs it. Even with a few apps running in the background
The free market, in its beautiful efficiency, leads to the intentional crippling of millions of state-of-the-art chips.
Try this thought experiment:
1. Based on market analysis, Intel decides it could sell a part with hyperthreading for $1K, and one without for $500.
2. Engineers start building both chips.
3. Because the chips themselves cost little to make ($45 and $50 for non-HT vs HT let's say?), and it costs $10M of engineering time to design each chip, engineers realize it would actually be far more efficient to design one chip with hyperthreading, and disable it for the lower-end SKU.
I'm curious which part you object to in this sequence. The sale price of chips is mostly amortizing very high R&D costs, not unit distribution costs. There are a variety of different types of customers to serve with different price points, while designing different chips is expensive, so it ends up being logical to make one chip with different features enabled.
Software and web services are the ultimate expression of this kind of economics. It costs almost 0 to serve an additional customer, but a lot of R&D and operations to build it.
Would you call a web service "intentionally damaged" when they don't give you all features for the same price (or for free?)
If the HT chip costs almost the same as the non HT chip then maybe they should just be selling HT chips for $500 instead, that would probably be what would happen if there was actually competition in the x86 market.
The number of transistors involved in HT are quite small. There are probably very few cores where one thread works and the other doesn't, compared to the number of CPUs that work or don't work because their cache or something more fundamental is screwed up by a defect.
I wouldn't be surprised if they don't even test for one thread working and the other not, and the "defeaturing" is a separate step after yield binning.
Do I ever use more than 2 cores? Rarely.
Oh actually it has 4 cores, guess that changed with the newer chips.
Yup. That changed with the 8000-series. The current i3 chips are more like the previous generation's i5.
It's pretty nice because it means i3's are suitable for gaming.
Core i9-9900K with 8 cores and 16 threads
Core i7-9700K with 6 cores and 12 threads
Core i5-9600K with 6 cores and 6 threads
[1] https://hothardware.com/news/core-i9-9900k-coffee-lake-cpu-i...
So say your CPU has 100 adders, and when you resolve all your dependencies for incoming instructions, you can only use 60 of those adders when running instructions in parallel (out of order), so the HT/logical core uses the other 40 adders for another thread (so the HT cores get a lower priority than the standard cores; hence why operating system that are HT aware can be more efficient by scheduling lower priority threads on the HT cores).
Is that correct or am I way off? (HT wasn't a thing when I took architecture class. I had a dual AthlonXP back then, where I had two physically separate processors).
But the biggest win is when HT can fill pipelines holes caused by a single thread waiting for some long latency operation (almost invariably memory accesses). Normally out of order execution can fill these holes by extracting parallelism from a single thread of execution, but that's not always possible when multiple high or unpredictable latency operations are chained together.
What this means is that hyperthreading does not really work so well on sustained, homogenuous workloads. For example doing very heavy computation on 8 threads of a 4-core CPU with 8 hardware threads, can actually reduce performance, because all threads will be contending for the same functional blocks of the CPU.
When I got my i7-3770K (4 cores, 8 threads) years ago, I was into POV-Ray rendering. I did a test using only 1, 2, 4, and 8 threads.
As expected, 2 threads was double the speed of 1, 4 threads was double the speed of 2, but 8 threads was only about 15% faster than 4.
Thinking about it now, I wonder what power consumption looked like.
what would distinguish this description from, say, compression and compiling? based on the experience on my machine (which of course is limited), HT does give a boost in those two cases, which could be classified as homogeneous and sustained.
> what would distinguish this description from, say, compression and compiling?
The real difference is between integer and floating point workloads.
Integer workloads typically have lots of branch mis-predicts and cache misses which cause pipeline stalls. Any time you've got pipeline stalls, hyper-threading will be great.
Floating Point workloads typically are just huge amounts of number crunching with very simple access patterns. There are few branches and memory accesses typically have a consistent stride and so are pre-fetch friendly. Typically, this kind of workload is memory bandwidth limited because your CPU can run full bore without any bubbles in the pipeline. Hyper-threading isn't much use in this case: if any functional units are going unused, it's only because the CPU can't vacuum data up fast enough. This is one of the big reasons GPUs are so popular for doing FP workloads: they have a ton of functional units and they are paired with very high speed on-board memory.
In the grandparent's defense, floating point workloads tend to be sustained and homogenous. They are, however, not the only kind of sustained, homogenous workloads, as you so accurately point out.
Compression would be closer to homogeneous. But something REALLY homogeneous is like, Prime95. You're only hitting SSE and/or AVX instructions over-and-over again. All threads try to only use AVX instructions, so hyperthreading doesn't help too much.
We have exactly this scenario on a server we're running (non-virtualised) and, if it didn't involve rebooting the machine, I'd already have run a test to benchmark whether we get better performance with or without hyperthreading. Unfortunately our test environments aren't similar enough to be representative in terms of testing yet, but this is going to change in the next few weeks so I'll finally be able to run my test without needing to take down the production box.
Example: https://www.golinuxhub.com/2018/01/how-to-disable-or-enable-...
Cycle 0: Thread 1 loads
Cycle 1: Thread 1 executes & thread 2 loads
Cycle 2: Thread 1 loads & thread 2 executes
Cycle 3: etc.
The advantage of hyperthreading is that there is little or no context switching overhead. At least relative to traditional hardware threads. With a real CPU there are more steps in the pipeline and each step takes a variable amount of time. Later CPU's take advantage of those available cycles with parallel, speculative and out of order execution...but those are generally abstractions over a single thread.SIMD is one thread doing one instruction on multiple pieces of data at the same time.
SIMD can give higher throughout from the CPU, and you must organize your data types to use SIMD.
Symmetric multithreading is where software takes advantage of multiple logical hardware threads to do multiple pieces of work per clock cycle.
SMT can use HyperThreads and/or multiple physical cores, and/or multiple physical CPUs with one or more logical hardware threads each.
This would be quite the opposite of Intel's actual SMT implementation, which aims to keep all the parallel execution resources of an OoO core fed and busy.
There have been some machines that do what you describe too (eg Tera/Cray MTA, many GPUs, Sun's Niagara), to combat memory latency and reduce the need for cache. Those machines have a big thread count since they want to have a lot of outstanding memory operations in flight. You will notice that these machines are not called SMT, since the S stands for "Simultaneous".
I don't think that's quite true. Rather, the two "virtual cores" are mostly independent, but share some backend resources. For example, modern Intel CPU's can sustain execution of 4 instructions per cycle, and with hyperthreading, two of these instructions can be from one thread and two from the other. Or one and three, or four and zero. Both threads truly do execute at the same time, it's just that the competition for resources means it sometimes takes longer for them to execute on a shared core than on separate physical cores.
https://www.youtube.com/watch?v=Xf0VuRG7MN4 (AMD results start at about one minute mark)
The issue with Spectre is that it is a new "buffer overflow". You can't "fix" a buffer overflow through hardware alone. You need software + hardware... and at best, you only get mitigations.
And within the next few months, some researcher is going to come up with a new Spectre-based attack that current mitigations won't work on. Its a bit annoying. Just sit tight and stay up to date on Spectre, its a moving target.
That's only relevant if you only plan on running one application at a time.
So their goal would be to limit meltdown to the flagship i9 series? Seems like a strange plan.
However, that is more for Spectre, which abuses the reads done from foreign code to observe the effect they have on the cache. That foreign code can be the kernel, another process, the hypervisor, etc. running on the same core. Meltdown can read the entire comments of memory without needing foreign code and therefore SMT doesn't matter for it.
Of course this doesn't mean its brand/sales won't be hurt further due to this move.
That said it would simply mean that the i7 would become the new i5 and also fill the same price range quite likely.
They might run the i5, 7 and 9 in the mainstream line now and bring the i3 to where the Pentium line is currently is however I still think that this is quite likely a configuration issue than a new segmentation since the i9 is their HEDT platform now.
choice is bad! :D
1. How much can I afford?
2. In that budget, do I need maximum instructions per second on one core or maximum cores?
If maximum cores, AMD is over there.
Remember that architectural registers (such as RAX or EAX) are "fake" and remapped to the real, physical, microcode registers. Code like "xor eax, eax" is translated into "physical-register #51 = 0" with regards to uops, and doesn't even use any execution units on modern Skylake or Zen processors!
Hyperthreads only need to contain one more remapping of physical registers to architectural registers.
* Parts with broken HT can be shipped as fully working non-HT parts.
* High-leakage parts can be shipped as they now pass the power screening.
It's unlikely that they ever get parts with broken HT on an otherwise salvageable core. Too much of that functionality is simply partitioning existing resources in half. There aren't that many transistors that are simply HT overhead that the core can do without when HT is not in use.