I don't understand why parallelization of a kernel task is a good idea. Zeroing a page shouldn't take that long. Especially on modern processors. Also what else is a heavyweight task besides zeroing a page?
I always thought x86 could benefit from block memmove/memset hardware external from the CPU.