It's an interesting idea, mostly applicable to long running processes (e.g,. a server) where the concept of "idle time" applies.
For something like a process that starts up, does its job as quickly as possible and then stops, it is less obvious how to do it: you could certainly start another thread to do this, but if you are willing to add threading you are probably better off just trying to parallelize the underlying work, rather than doing zero async.
The main downsides to this approach:
1) It is cache unfriendly: if you zero the pages synchronously immediately prior to use, you bring them into cache during the zeroing and then they are hot when you use them on the same thread. For idle zeroing this link is lost so you are likely to bring them into cache twice: once during the zeroing and once during the subsequent use.
2) On NUMA systems, doing the idling on some thread/CPU distinct from where you will use it is an anti-pattern: with the usual "first touch" allocation, the memory will be allocated on the node doing the zeroing, potentially different from the node doing the using - which could be an arbitrarily large penalty over time (as you keep using the remote memory for an arbitrarily long time)
Note that (1) is a common confounding factor when measuring the cost of zeroing. E.g., if you have something like:
zero(mem)
do_work(mem)
You might measure the zeroing taking 500 (time units) and the work taking 800, total of 1300. Lets say you find a clever way to totally avoid the zeroing (e.g., maybe do_work didn't actually need zeroed memory): you expect only the 800 time units of work to remain - but actually measure 1200: the zeroing more than just zeroing: it was bringing in all the memory into cache and suffering all the cache, TLB misses, etc. When you remove it, that work just moves over to the subsequent work.
In some cases do_work alone can even be slower than the sum of zeroing and do_work (e.g, 1,500 time units) - because the interleaving of the work with cache misses can be less efficient in an MLP sense (because the work clogs up the OoO buffers so fewer memory requests fit in the window).