So does that mean that further major optimizations are basically impossible?
You don't always need the operation finished right away. Maybe REP MOVS could return a sort of a handle that you then wait for when you actually need to use the destination, and the CPU can keep chugging along in the background. Like when you submit commands to the GPU or how you can wait at memory fences for other threads in a warp to complete.