Anyone knows what kind of issues having "The frontend, FPU, and L2 cache [] shared by two threads" causes compared to 2 real cores / 2 hyper threads on 1 core / 4 hyper threads on 2 cores ?
Anyone knows what kind of issues having "The frontend, FPU, and L2 cache [] shared by two threads" causes compared to 2 real cores / 2 hyper threads on 1 core / 4 hyper threads on 2 cores ?
I wouldn't say there are issues with sharing those components. More that the non-shared parts were too small, meaning Bulldozer came up quite short in single threaded performance.
Long story short, running two threads with only enough hardware to run just one at full speed means any thread that tries to run at full speed will not be able to.
> Each decoder can handle four instructions per clock cycle. The Bulldozer, Piledriver and Excavator have one decoder in each unit, which is shared between two cores. When both cores are active, the decoders serve each core every second clock cycle, so that the maximum decode rate is two instructions per clock cycle per core. Instructions that belong to different cores cannot be decoded in the same clock cycle. The decode rate is four instructions per clock cycle when only one thread is running in each execution unit.
...
> On Bulldozer, Piledriver and Excavator, the shared decode unit can handle four instructions per clock cycle. It is alternating between the two threads so that each thread gets up to four instructions every second clock cycle, or two instructions per clock cycle on average. This is a serious bottleneck because the rest of the pipeline can handle up to four instructions per clock.
> The situation gets even worse for instructions that generate more than one macro-op each. All instructions that generate more than two macro-ops are handled with microcode. The microcode sequencer blocks the decoders for several clock cycles so that the other thread is stalled in the meantime.