On newer Intel CPUs, REP STOSB is highly optimized, but you still have to do the writes.
On newer Intel CPUs, REP STOSB is highly optimized, but you still have to do the writes.
Problem is: how does the CPU know it doesn't have to do that read? REP STOSB may, but it isn't trivial to implement, as neither start nor end of the range to be processed need to lie on a cache line boundary.
PowerPC has the dcbz and dcbzl instructions which zero a cache line without reading the to-be-zeroed data (dcbzl was invented because real-world code assumed dcbz always zeroed 32 bytes. See http://lists.apple.com/archives/darwin-drivers/2005/Apr/msg0...)
The cache can be bypassed when doing a lot of writes. See section "Bypassing the cache" of [0].
That article says that memset() already uses cache bypassing instructions for large blocks.