Thoughts on Skymont Slides
chipsandcheese.com
chipsandcheese.com
> This scheme appeared with Tremont, where it could only switch decoders at taken branch boundarie and Gracemont improves load balancing between the two decode clusters by automatically switching, instead of relying on taken branches. That helps performance in very long unrolled loops, which could get stuck on one decode cluster on Tremont. When we tested with longer loop lengths, we didn’t see any drop off in Gracemont’s instruction throughput:
- https://chipsandcheese.com/2021/12/21/gracemont-revenge-of-t...
Also back in those days 32bit x86 with it's anemic 8 GP registers (and stack-based-fpu) it was hard to avoid any dependencies w/o trying to calculate memory positions, nowadays we have 64bit x86 with 16 GP registers and separated FPU XMM registers (compilers can leave much more in memory and registers are probably far easier to do dependency resolution on).
Oh and a theory could be that JS,Java,C#,Rust,etc code with lots of non-taken safety barriers(out of bounds, type-tests,etc) has probably been analyzed and been found to be a prime candidate for parallell issue.
3x3 clustered decoder, 8 ALU units, and 4x128 vector units. This really is a quiet exciting architecture.
How did they do that at 20 nm instead of 5nm ? Simpler cores.
Trying to improve single-threaded performance is really just adding epicycles and we need to just move onto more cores. Most of the time your CPU is just waiting, otherwise there wouldn't be all this focus on boost frequencies.
Slow it down it's reliably fed, get rid of the power throttling, because then it's not needed. Less surface area for bugs or attacks. An army of ants is a powerful thing.
Also, FWIW, Xeon Phi hit 244 threads in 2012 and 256 threads in 2016, although it used 4 threads/core.
Sparc development could focus on servers and many of those tasks were easy to parallelize so of course they went for cores once they couldn't keep up on the frequency race (esp when selling expensive machines).
doesn't this mean more single thread performance is wanted/needed?