Because DRAM is (relatively) so slow, avoiding the _reads_ is a big gain. If you zero a large range of memory the naive way, the first write to each cache line _reads_ a cache line of data that likely will be overwritten by zeroes.
Problem is: how does the CPU know it doesn't have to do that read? REP STOSB may, but it isn't trivial to implement, as neither start nor end of the range to be processed need to lie on a cache line boundary.
PowerPC has the dcbz and dcbzl instructions which zero a cache line without reading the to-be-zeroed data (dcbzl was invented because real-world code assumed dcbz always zeroed 32 bytes. See http://lists.apple.com/archives/darwin-drivers/2005/Apr/msg0...)