I've literally never heard of this. I know you can have multiple execution units for parallel instruction execution, but 'light' and 'heavy' - can you give some info or a link? TIA
I've literally never heard of this. I know you can have multiple execution units for parallel instruction execution, but 'light' and 'heavy' - can you give some info or a link? TIA
https://www.anandtech.com/show/13699/intel-architecture-day-...
I believe this shows differences between FP and Integer units. In order to achieve a certain performance goal you don't necessarily need another integer divider when you want a new adder/multiplier. So you add a slimmer unit instead.
I listed this in my original comment, because this is a giant can of worms for the compiler and decision-maker on where to execute what.
With respect, I think you're misunderstanding. I thought you meant light/heavy versions of eg. adders, for some definition of light and heavy addition.
I'm not an expert but... CPUs will put in extra execution units according to need (will typical code get faster with an extra X?) and cost.
Shifters are typically very often used, and are simple. So are adders, though more complex. IIRC recent intel x64 will have several of of each[0]. Multipliers are less cheap so they have fewer (and often you can turn them into adds in certain cases such as progressive array lookups). Division is slow and very expensive in transistors, so they have 1 (division can often be turned into reciprocal multiplication anyway). Sqrt is even worse.
And to repeat, I'm no expert and any corrections welcome.
[0] <https://en.wikichip.org/wiki/intel/microarchitectures/coffee... If I'm reading this right, 2 shifters (2? I suppose they are fast so they are available soon after), 4 adders, 1 mult and 1 divider.
Further thinking suggested there'd be a ton of wires doing this, and perhaps it's the wiring that's taking up the silicon?
Do ports (as in the picture) have independent pipelines, or do they execute certain pipeline stages of a big pipeline? I suppose, either way you can't issue to the same port in the same cycle.
This paper sheds some light on how instructions are divied up between units on NVIDIA GPUs. http://www.stuffedcow.net/files/gpuarch-ispass2010.pdf Table IV. Notice that fp32 mul is in the SFU and SP, while others are not.
<https://en.wikipedia.org/wiki/Re-order_buffer>
Come to think of it, I don't know how the ports are used. I am entirely unqualified to answer this question :)