On a Xeon v3 core, SPECint averages below 2 instructions per cycle: https://www.researchgate.net/publication/322745869_A_Workloa.... How does Apple beat Intel on branch integer benchmarks like 403.gcc by a factor of two per clock?
On a Xeon v3 core, SPECint averages below 2 instructions per cycle: https://www.researchgate.net/publication/322745869_A_Workloa.... How does Apple beat Intel on branch integer benchmarks like 403.gcc by a factor of two per clock?
Mostly I'm curious about how complete the bypass network is on their functional units and if execution is clustered like the POWER8. The width doubling in the A series does remind me of the POWER 7 to 8 transition.
Renaming is also apparently a major constraint on design width in many cases but I'm not so familiar with that.
What does "FO4" stand for here? Googling it yields "Fallout 4", which definitely isn't right, and I'm not sure what other keywords to tack on to get the right result.
> Fan-out of 4 is a process-independent delay metric in CMOS tech.
The wrap-up slide at the very end has a bit of a rundown.
http://people.duke.edu/~bcl15/teachdir/ece590_fall14/Present...
You can check that part at least. Isn't it LLVM?