Intel's Redwood Cove: Baby Steps Are Still Steps
chipsandcheese.com
chipsandcheese.com
Same for branch delays slots.
Were there a lot of idle/no-op stages for simple instructions? Or did most instructions, even simple ones, actually get some useful work done for each of those 32ish stages?
Mostly the stages you expect from an out-of-order core all had their own pipeline slot in order to go super fast, and some were split into two stages. For example, you have two stages to calculate the next instruction (branch prediction), two stages to fetch, a stage to redrive a value from the memory bus, etc. From there, the ALUs could only do a 16-bit add in one cycle, so even an addition took at least 3 cycles. The later Pentium 4's took this subdivision even further with the goal of reaching 4-5 GHz clocks in 2005.
The whole chip was hand-laid-out, so it was possible to put pipeline registers in places that would have been weird to do in Verilog.
When you encounter a biased-taken branch that isn't currently being tracked by the machine, you're always condemned to pay the cost of a misprediction because the default prediction is "not-taken". The hinting here is supposed to indicate "when encountering this branch, don't predict not-taken by default."
[1]: https://gcc.gnu.org/onlinedocs/gcc/Other-Builtins.html#index...
Now the execution of both sides of a branch is basis for the 'Specter' side-channel attack. Access restricted data in the false branch, data access still happen even though it was restricted.
That is essentially impossible to prevent and disabling it killed a lot of cpu performance, branch hints might be the next best thing.
Also, I don’t think executing both sides of a branch ever took off on any mainstream CPUs (unless Itanium counts). It wastes power to spend half your execution units on things that won’t be committed, why not use them on the other hyperthread instead?
That part I am sure off. I will double check with a friend of mine of of curiosity, but one thing to note is that the execution units are processing branches up to the point when branch is evaluated, then the false path is dropped.
Back then the speed was trumping the power draw. I am not sure what are the priorities today.
In terms of hyperthread, i don't think you can safely execute instructions of both siblings due to possible shared cache mem clashes. But I am guessing now. Its been a while since I have been working that low level to remember the details.
Modern CPUs speculate hundreds of instructions ahead, and with just a dozen branches you can have a few thousand different paths. It makes more sense to speculate down one path with very high accuracy.
I think a lot of folks get mixed up with GPU and/or SIMD architecture, where you execute both sides of the branch out of necessity: Some of the lanes need to go one way and some the other, so you have to do both.
not to mention that you burn a lot more power.
By all accounts of Spectre I've seen, the processor never tries to speculate both sides of the branch. Instead, it always predicts one side of the branch, and speculatively executes that side only: when it later turns out to have been a misprediction, it rolls back the speculative execution, and proceeds to execute the other side from scratch. Both sides are executed at some point, but not simultaneously.
The only purpose of the branch predictor, which has become one of the biggest and most important parts of any modern CPU core, is to execute only one of the two branches, hoping that you have guessed the right one.
On a CPU that executes both branches, there is no reason for a branch predictor to exist.
The equivalent of executing both branches is obtained by replacing the conditional branches with conditional move or conditional select instructions, which are used after both values corresponding to the two alternatives have been computed. The use of conditional move/select instructions is justified only when the direction of the conditional branch would have been quasi-random, so the branch predictor would have failed to predict it.
Executing both branches has an exponential cost in the number of conditional branches that are speculated ahead, so it is neither feasible nor desirable, as it would greatly increase both the power consumption and the die area for a given performance.
https://en.cppreference.com/w/cpp/language/attributes/likely
One of the things that the BOLT post-link optimizer does is rearrange basic blocks to reduce taken forward branches, based on the profiles.
I have absolutely no idea how important that it. But there exist that bit of extra complication on the real world.
That is, the stuff that will go into a retry of something likely has far more setup than your typical hot loop.
It's a fun play on the idea of a productivity enhancer. AI - the extra fingers or hand - does well to spare me effort, I just think productivity may not be the right way to look at things
Huh? Who?
(But also, you must go watch “The Princess Bride” tonight to catch the million cultural references you’ve been missing.)
The article points out specific use of specific techniques to probe cache and other efficiencies, which is valuable.
Now, everyone gets five fingers and a thumb, as requested.