glibc's is all in assembly full of SIMD instructions, which seem very much magical...
https://sourceware.org/git/?p=glibc.git;a=tree;f=sysdeps/x86...
http://stackoverflow.com/questions/8858778/why-are-complicat...
Also, I wonder if zeroing large chunks of memory would be faster to do in kernel space using real addresses. You can avoid the multiple real memory lookups involved in a single virtual write.
(Of course, we already avoid those often, but it could be useful to avoid entirely. Not sure what the tradeoffs are here)
(The kernel's linear mapping of physical memory can take advantage of huge pages though, which means that there might be one or two less levels of page tables involved with those addresses). TLB misses aren't significant if you're bulk-writing to a block of memory anyway, you'll max out the bandwidth of the memory without that being an issue.