REP MOVS is still extremely fast --- and also small. It'll copy cacheline-sized chunks if the size is large enough.
LOOP is a bit of a weird case. I've seen it benchmark both slower and faster than dec/jnz depending on the surrounding instructions.
LOOP is a bit of a weird case. I've seen it benchmark both slower and faster than dec/jnz depending on the surrounding instructions.
now, does anyone know why?
My guess is virtual stack pointer update prediction latency.
To expand on that, Intel's CPUs have had for a long time a separate piece of hardware dedicated to a "virtual" stack which speeds up push/pop instructions. If pushes and pops are not mismatched, then all stack operations can stay entirely within that and there's no need to update the "real" stack pointer nor stack entries upon leaving the loop.