1- using the "INC mem" form is not a fused instruction. Intel has macro op fusion where two macro ops are send to be decoder as a unit like cmp+jmp to be decoded into uops more efficiently. This isn't that.
2- micro op fusion is when after decoding from macro to micro op, two micro ops that are adjacent are retired together, just as "ADD reg mem" goes to the micro ops "LOAD reg2 mem" and "ADD reg reg2".
Or in his example:
INC mem ->
LOAD reg mem
INC reg
STORE mem reg
and the INC and STORE will be fused and retired together (not the LOAD and INC), IIRC.
3- he's hitting the loop stream detector so decoding is bypassed anyways. His 3 macro op version is hitting the uop cache, same as tbe 1 macro op version.
4- he's trying to say both the "good" and "bad" will execute the same, but i dont think this is impacted by the uop fusion. And nothing in his counters shows any difference.
The single direct memory form of INC cuts down on register and L1 instruct cache pressure, neither of which this benchmark will pick up.
I'm a software guy though, so corrections are definitely encouraged.