The second one is that mmap prevents you from a tight control of memory in various pipelining scenarios when you want to have your memory overhead to be constant and not sizeof(data).
But the alternative we have suffers from the same issue, doesn't it?
malloc is implemented in terms of sbrk and mmap/munmap, both of which are guarded by locks. And malloc will also have its own user-space locks on top of the ones from the kernel. I checked glibc and jemalloc implementation.
> mmap prevents you from a tight control of memory ... when you want to have your memory overhead to be constant and not sizeof(data).
I am not sure I understand this argument. It's trivial to implement the user-space custom MMAP allocator that gives you no less control over the memory than other "types" of allocators. Actually, it can give you more power because you can avoid the locks from user-space malloc implementations if you know that the use-case you're crafting it for is single-threaded. I used this technique in the past for short-living objects and in certain workloads it improved the performance by a factor.
Nobody is arguing that you shouldn’t call mmap to get yourself a few anonymous virtual pages to do your work in.
However, I could imagine that it is possible for contention in kernel-space mmap locks could be artificially relieved because of the user-space locks in malloc.
Currently I don't see how TLB shootdowns would be specific to mmap only but not to malloc but I may find some time to read the paper.