The operating system is free to load and cache the file into memory in whatever way it wants, so for large memory mapped files it'll often try to use all available memory to cache as much as possible. If another program needs memory, the operating system will simply lower the amount of memory available to the memory mapped file for caching. This is extremely useful for databases, since it greatly simplifies both how to persist the data along with how to load and cache the persisted data.
This all comes with a big caveat however. The less memory you have, the more file accesses occur (similar to your pagefile when you're thrashing), which can dramatically slow down your memory operations.
tldr; it lets you designate a file to use as a region of memory.
It saves a copy on read, in exchange for a bit more overhead configuring the processes address space up front.
> "The fundamental operation of mmap is to add new entries to the page table of a process, and the precise properties of those entries are heavily dependent on what the arguments to mmap are.
> When you mmap a regular file, you're essentially adding an entry to the page table that shares the data with the kernel's filesystem cache."
Now, understanding this requires some understanding of an OS's virtual memory function, and what "page tables" are. Those are what the OS use to track a process' memory, whose granularity is in "pages" (historically 4kB on Linux, though others use larger granularity such as 16kB).
mmap() has a a lot of flags [2] that affect the properties of those mapped pages. It is the swiss-army knife of memory management on Linux.
Some of those properties allow you to share memory with other processes (MAP_SHARED | MAP_ANONYMOUS), or just allocate memory (MAP_ANONYMOUS) or, by default, map an (open) file specified by the `fd` argument.
(fun fact! On linux, when you malloc(), you don't actually get memory, just an IOU from the kernel. Only when you access that memory, ie accessing those memory pages, does the kernel actually make the effort of allocating you that memory.)
It seems the intuitive way to implement read if there's a complex memory subsystem that does all these great things anyways.
The OS somehow detecting that argument to read(2) is page aligned block of uncommited meory and transparently creating MAP_PRIVATE mapping from that seems like nifty idea, but userspace tends to not have page-aligned buffers that often (see the aforementioned 0x10 offset for one reason why) and applications that intentionally allocates IO buffers with some particular alignment do so in order to bypass the block cache for particular IO patterns and are exactly the kinds of applications that tend to use mmap(2) for IO where that does not matter.
Hmm. You're right about glibc, but macOS for instance reliably returns page-aligned pointers for large allocations. I wonder how Windows behaves.
Well, first of all, it's not: malloc may have previously stored its internal structures in that memory so it may be dirty. Second, AFAIK if you ask malloc() for a gigabyte, it'll internally call mmap() on an anonymous file and will give you its result.
But from what I understand there may be nothing in POSIX forbidding the kernel to do this optimization, just kernel authors decided against it (which is understandable).
I can think of one edge case not mentioned by others, though.
One of the main benefits of mmap is that the data can be paged out while not used. But when you try to page the data back in, the attempt may fail, or it may return different data than was originally there. This is particularly likely with network filesystems or external storage. With mmap, these situations result in, respectively, the process receiving SIGBUS, or the contents of memory changing out from under it. The former is usually a crash (unless the program installs a signal handler), and the latter can be worse than a crash, if either program logic or compiler optimizations make assumptions about memory staying consistent when not written to. A program using mmap is opting in to that risk. A program using malloc and read is not.
That said, it may be reasonable to trust that the system’s main disk(s) (however you define that) won’t misbehave this way, and limit the optimization to files mapped from them. After all, if swap is enabled, the same potential risks exist for any memory that’s swapped out. Even without swap, if the root partition starts failing, the system is not going to stay up for long.
Or, the optimization could be enabled for other disks, but with the caveat that the pages would be swapped to the swap file/partition rather than paged out normally.
Still, the current problem with mmap is that it's buggy on some operating systems according to https://www.sqlite.org/mmap.html, so I guess it will cause more frustrations until AGI takes over the world :)
When a read call completes, you know you've successfully read the data from the underlying device. If you have swap disabled you know you can read and write from that memory quickly.
With mmap'd data, a memory read/write can trigger a page fault and I/O, and if you encounter an error like a network filesystem not responding in time, the program sees a segfault.
`mmap` signals the kernel that you want to use the data someday. You `mmap` it, it automatically handles fetching and throwing unused bytes away.
`read` signals the kernel that you want to do something with the data now.
You `read` it, you need it in the memory now.
> why the kernel can't see that it's unused allocated memory
Only a in-process garbage collector would really know what part of your memory is unused. The kernel can't really know for sure.
https://manybutfinite.com/post/anatomy-of-a-program-in-memor... [1]
You can't really understand mmap/sbrk without understanding virtual memory and process space layout.
[1] Images are broken. Just open https://static.duartes.org/img/blogPosts/kernelUserMemorySpl... and go to advanced and "Proceed to static.duartes.org" to workaround their https issues. The duartes.org host is owned by the blog author. Refresh the blog article and images should load now.
The difference is between reading a file and memory mapping (mmap'ing) it.
If you read a 1TB file into memory you use 1TB of disk and 1TB of physical memory. If you then access that data it's as fast as RAM because that's where it is.
If you mmap a 1TB file you use 1TB of disk and 1TB of virtual memory. If you then access that data it may be mapped into virtual memory but not actually be in RAM. This triggers a page fault, at which point the correct page is loaded from disk to physical memory, and handed back to you.
The key observation is that the amount of physical memory occupied by the mmap'd file is much smaller than the entirety of the file unless you access almost all of it.
If your filesystem supports holes, this can also be useful for writes: it's possible to map files vastly larger than physical disk space, but so long as the actual number of places written to is quite small you won't run out.
The combination is very useful for datastructures because you basically don't have to care about data being extremely sparse until that data is also getting quite large, which means you can use cheap/fast approaches to indexing, etc.