Actually, I think it does: you cannot be using the core while it's doing the memset or memcpy, so it's technically not what I'm describing.
Even if it did: a cross-industry reference implementation would go a long way into making this a reality.
I'm willing to bet that within 5 years we'll see a CPU that effectively embeds a DMA engine used through this instruction. The way I'd implement it is a small FSM in the LLC that does the bulk copy while the CPU keeps running, while maintaining a list of addresses reads/writes to avoid (i.e. stall on) until the memcpy is finished.