Ice Lake Store Elimination
travisdowns.github.io
travisdowns.github.io
You can find the previous article at [1], which goes over the basics and the original finding on Skylake. This new article focuses on Ice Lake but probably mostly makes sense after reading the original. HN discussion of the first part is at [2].
---
[1] https://travisdowns.github.io/blog/2020/05/13/intel-zero-opt...
Let me whip something up.
The book that got closest for me was CS:APP, but haven't found anything beyond that. It was also very generic. I'm looking for specifically high-perf optimization techniques for x86-64
Any specific books, websites (other than Agner's) or even other blogs?
https://github.com/mfleming/performance-resources/blob/maste...
How To Write Fast Numerical Code: A Small Introduction - Srinivas Chellappa, Franz Franchetti and Markus Püschel
http://spiral.ece.cmu.edu:8080/pub-spiral/abstract.jsp?id=10...
https://github.com/travisdowns/travisdowns.github.io/issues/...
Screenshot appreciated if you have a second. Trying to find a way to run WebKit on Linux...
Any other rendering issues let me know.
So, its pretty obvious they are choking off the single core bandwidth to keep a single core from starving the others in the machine. This 100% makes sense for server applications, but for workstation and desktop usage its completely crazy that a single core has 50-60 GB/sec L3 bandwidth, but can only utilize 10-20% of the available ram bandwidth to flush the L3 even when the other cores in the machine are idle.
The ARM machine literally has somewhere between 2-4x the single core memory bandwidth, despite being in a configuration with more actual cores than intel offers.
That said, in ICL there has been a nice jump.
Graviton results to RAM certainly are nice. I've heard that in this type of store workload they can automatically use something like NT stores, avoiding the RFO (which would normally cut bandwidth in half).
The Intel chips do get much better numbers if you use NT stores... but I didn't go into that since it wasn't the point of this article.
A decade+ ago, Andy Glew expounded on the mistakes they made for the P6 when it came to the rep prefix. Apparently, they have finally fixed the startup times for it in icelake. But it continues to be a squandered opportunity because in theory they could have hidden a lot of the generational/microarch messiness behind the rep mov/sto sequences rather than requiring software updates every-time they updated the microarch like a RISC machine. AKA rep should be a lot harder to beat on any given machine.
The nontemporal case is another one of these instances where the rep sequence could signal early on that the operation is going to be more efficient with a NT store. Reserving actual software defined NT stores for the cases where software is absolutely sure it won't fetch the data in the near future.
There are a few other cases where high level microcoded instruction like this are an advantage because it allows a common piece of code to be implemented in differing ways depending on the actual product. Reference, Alpha PAL code.
What really surprised us was how much faster Excel went, it wasn't really our target, turns out that every time it updated the screen it redrew the background 5-6 times - drawing white over white - the actual content (the black pixels) were just noise in the graphics numbers
I'm not sure what you mean by the 512 downlock: nothing I describe in the article has much to do with a downclock, except for the L1 effect at the beginning, and this was honestly fairly specific to the original test structure, which alternated periods of spinning on a timer with running a short test interval. It wasn't actually a downclock either (this CPU does not downclock for AVX-512 at 3.5 GHz: it has very little downclocking at all): rather it's dispatch throttling: the CPU still runs at full speed, but instructions are prevented from dispatching every cycle, effectively slowing the throughput, until the voltage can adjust to the heavier instructions.
"Store ports" are only a concept that apply to the core itself, not the caches. Ports are entry points for execution of a certain type of instruction. Older CPUs had one store port (p4) and Ice Lake has two (p4 and p9), so two stores can execute in a single cycle. So these ports only benefit the core and don't interact with the L2 or L3 directly, which operate on a cache line basis.
Caches also have something called ports, e.g., a cache might have 2 read and 1 write ports, meaning 2 reads and 1 write per cycle, but there isn't any indication these have changed in Ice Lake.