The Case of the Missing Increment
computerenhance.com
computerenhance.com
While the loop he's testing is a useless bit of code that does nothing the optimisation he's discovered may help speed things like scasb/stosb allowing portions of 2 unrolled copies to be processed per clock
Well, you can pick up Sapphire Rapids instances from your preferred cloud provider and avoid the sadness.
dunno about others
Normally it would be the either the programmer's or the compiler's job to unroll a loop and then reduce dependency chain lengths.
But its nice if the renamer can do that as well.
Presumably intel have real-world data that suggest that significant real workloads can profit from this.
I wonder whether that points to specific software issues, like hypothetically "oh yeah, openjdk8 hotspot was a little too timid at loop unrolling. It won't get that JIT improvement backported, but our customers will use java8 forever. Better fix that in silicon".
I don't see a reason why this should be the case, since the high bits of the result would simply be cleared, and it's a common size optimization to use 32 bit operations.
Maybe https://news.ycombinator.com/item?id=41706743 is correct, and this is mainly intended for address increments generated by microcode?
I looked it up, "gress" comes from "gradi" in Latin which directly translates to "walk". More specifically: con(pro) + gradi -> congredi (verb) -> congressus (noun)
Edit: Knowing this, "gradient" has an interesting flavour :)
Edit: It looks like the path is more indirect for "gradient"
"gradi" (walk) -> "gradus" (step) -> "grade" (french influence) + "salient" -> "gradient". I like that in Latin "walk" is "to step", or perhaps "step" is "the unit of walking"? "A walking"? Etymology is fun!
Consider the verb "to pace", and the corresponding noun "pace": the analogy is almost perfect. Of course, Latin also had other words for going places.
https://stackoverflow.com/a/58146426
Also in the bad old days SMM would interfere on some CPUs.
I guess immediate addressing mode addition is a good choice to execute at rename / allocation stage, as it's common, relatively simple and can't generate exceptions.
Well, except for the fact that you need to read from a register before adding the immediate displacement to it. You'd have to know the physical register and do the read very early (before renaming), or predict the value!
Caveat is just that [presumably] the source and destination registers have to be matching (since `lea rax, [rax+imm]` is just `add rax, imm`).
Similarly, 'lea r64, [r64+8]' (imm8) and 'lea r64, [r64+128]' (imm32) and 'add r64, 2' (imm8); but not 'add r64, 0x1000000' (imm32).
[0]: https://uops.info/html-lat/ADL-P/INC_R64-Measurements.html
[1]: https://uops.info/html-tp/ADL-P/INC_R64-Measurements.html