Avoiding Instruction Cache Misses
pdziepak.github.io
pdziepak.github.io
The performance of software is influenced by the workings of tens of units (+local caches) connected (non-linearly) with fifos, ooo buffers, replay mechanism. There are multiple versions of the same unit to save chip space (like light and heavy Integer/FP). Units/domains have different clock speeds. And from generation to generation port connections and instructions between pipelines can be reshuffled.
I suppose what keeps performance changes relatively straightforward for developers is the set of benchmarks used to evaluate the hardware early on.
I appreciate the article focusing on the I-Cache, and the nice intro to decoding.
I would have preferred having an example, and improving something instead of abstractly talking about problems, effects of code and optimizations and possible workarounds.
Tangent: I wonder if we will be seeing specialized instructions sometime that span multiple units and multiple cores in order to reduce data-movement. Think matrix-matrix multiplication. The potential improvements for power and speed seem huge.
I've literally never heard of this. I know you can have multiple execution units for parallel instruction execution, but 'light' and 'heavy' - can you give some info or a link? TIA
https://www.anandtech.com/show/13699/intel-architecture-day-...
I believe this shows differences between FP and Integer units. In order to achieve a certain performance goal you don't necessarily need another integer divider when you want a new adder/multiplier. So you add a slimmer unit instead.
I listed this in my original comment, because this is a giant can of worms for the compiler and decision-maker on where to execute what.
With respect, I think you're misunderstanding. I thought you meant light/heavy versions of eg. adders, for some definition of light and heavy addition.
I'm not an expert but... CPUs will put in extra execution units according to need (will typical code get faster with an extra X?) and cost.
Shifters are typically very often used, and are simple. So are adders, though more complex. IIRC recent intel x64 will have several of of each[0]. Multipliers are less cheap so they have fewer (and often you can turn them into adds in certain cases such as progressive array lookups). Division is slow and very expensive in transistors, so they have 1 (division can often be turned into reciprocal multiplication anyway). Sqrt is even worse.
And to repeat, I'm no expert and any corrections welcome.
[0] <https://en.wikichip.org/wiki/intel/microarchitectures/coffee... If I'm reading this right, 2 shifters (2? I suppose they are fast so they are available soon after), 4 adders, 1 mult and 1 divider.
Further thinking suggested there'd be a ton of wires doing this, and perhaps it's the wiring that's taking up the silicon?
Do ports (as in the picture) have independent pipelines, or do they execute certain pipeline stages of a big pipeline? I suppose, either way you can't issue to the same port in the same cycle.
This paper sheds some light on how instructions are divied up between units on NVIDIA GPUs. http://www.stuffedcow.net/files/gpuarch-ispass2010.pdf Table IV. Notice that fp32 mul is in the SFU and SP, while others are not.
<https://en.wikipedia.org/wiki/Re-order_buffer>
Come to think of it, I don't know how the ports are used. I am entirely unqualified to answer this question :)
It is unfortunate that this is still so useful in practice, given what it implies about the magnitude of waste in typical software systems.
Obviously, abstractions are useful, but it can be very hard to dig out of wrong abstractions, when they've influenced the design of the whole project.
This can go a long way to improving user code...but of course it's no magic bullet.
I think you’re right about a codebase replacing parts of STL as necessary. One I see frequently is `std::array` being wrapped with iterators to be essentially a stack-allocated vector.
I'm know little about this area, so I'm curious.
perf record -e icache.ifdata_stall ./test-program-binary
followed by perf report
or, for fancy viewing, use hotspot from KDAB.This is on linux, btw, but it should work fine with wine.
I guess a naive compiler could have issues but every time I check something on compiler explorer, the modern big boys (GCC, LLVM, ICC?) are pretty shrewd (Especially if you optimise for size).
I wonder if modern compilers would have allowed the team to keep at it.
It's been a while, but if I recall correctly, the optimized, stripped exe's at the time were a bit under 100mb. >500mb with debug symbols, etc.