A queue of page faults (2014)
zeuxcg.org
zeuxcg.org
Same question for memcpy and stuff like that: shouldn't that be a CPU instruction instead of special code trying to detect which CPU and emit optimized instructions?
Is my model of these costs that far off?
Trouble is, it's still just a simple loop, and it's not that much faster than doing it the long way... and on some architectures it's actually slower. It all depends on how optimised the microcode it decodes into is. After all, it still has to set the address bus and issue a write for every location in the RAM. There are no shortcuts there.
Some architectures will let you do memcpys and memsets via the DMA engine, which can give a small performance boost, but these are usually embedded architectures. (I've used one on the MSP430.)
As you mention, sometimes it's still slower and depends on microcode. A dedicated memcpy instruction shouldn't have that issue right?
On newer Intel CPUs, REP STOSB is highly optimized, but you still have to do the writes.
Problem is: how does the CPU know it doesn't have to do that read? REP STOSB may, but it isn't trivial to implement, as neither start nor end of the range to be processed need to lie on a cache line boundary.
PowerPC has the dcbz and dcbzl instructions which zero a cache line without reading the to-be-zeroed data (dcbzl was invented because real-world code assumed dcbz always zeroed 32 bytes. See http://lists.apple.com/archives/darwin-drivers/2005/Apr/msg0...)
The cache can be bypassed when doing a lot of writes. See section "Bypassing the cache" of [0].
That article says that memset() already uses cache bypassing instructions for large blocks.
According to MSDN, even when you use MEM_COMMIT: "The function also guarantees that when the caller later initially accesses the memory, the contents will be zero. Actual physical pages are not allocated unless/until the virtual addresses are actually accessed."
https://msdn.microsoft.com/en-us/library/windows/desktop/aa3...
If you want better hardware support, you need to keep the memory hierarchy in mind, and recall that there are multiple coherent masters in any modern system. If we teach DRAM to zero a block on command, the CPU can't just ask DRAM to do that - the L1, L2 and L3 caches need to be kept coherent, and the other CPUs have their own L1 caches too. We could teach DDR to accept "write zeros" commands from L3, and L3 to accept them from L2, and L2 to accept them from L1. L1 is already accepting them from the CPUs that have memory-zeroing instructions. In fact I'm aware of one CPU design where that already exists in the L1<->L2 interface [2].
One could also take a different approach - if you assume the kernel wants to keep the memory on the free-list zero'd, then you could have another coherent master in the memory system to do that work for the kernel. When a page is de-allocated, the kernel hands it to the memory zero-er, when the zero-er is done, it fires an interrupt and the kernel accepts the page onto the free list. I'll bet there are a lot of DMA controllers that could be made to do this already, although I don't know if any of them could do it fast enough to be a net win given the extra interrupt activity.
[1]: Except on some embedded CPUs, which can manage caches more efficiently when using memory-zeroing instructions.
[2]: http://infocenter.arm.com/help/topic/com.arm.doc.100511_0401...