The performance of software is influenced by the workings of tens of units (+local caches) connected (non-linearly) with fifos, ooo buffers, replay mechanism. There are multiple versions of the same unit to save chip space (like light and heavy Integer/FP). Units/domains have different clock speeds. And from generation to generation port connections and instructions between pipelines can be reshuffled.
I suppose what keeps performance changes relatively straightforward for developers is the set of benchmarks used to evaluate the hardware early on.
I appreciate the article focusing on the I-Cache, and the nice intro to decoding.
I would have preferred having an example, and improving something instead of abstractly talking about problems, effects of code and optimizations and possible workarounds.
Tangent: I wonder if we will be seeing specialized instructions sometime that span multiple units and multiple cores in order to reduce data-movement. Think matrix-matrix multiplication. The potential improvements for power and speed seem huge.