The mmap pattern
john.freml.in
john.freml.in
The cure is surely worse than the disease!
The program state in memory at the time of abnormal termination is likely to be inconsistent, leading to an unusable file. The subset of that that happens to have been committed to disk is likely to be worse.
Contrast with the traditional approach of explicit serialisation / deserialisation for persistence:
* serialisation occurs when the program is in a known state. We can reason about what invariants hold at the point when state is sent to storage. We can take steps to confirm that data actually has been committed to persistent storage before we treat the operation as complete.
* deserialisation recreates program state in a clean environment from minimal data. This has the side effect of reinitialising state that may have caused the unexpected termination. Coupled with the ability to reason about program state at the point of serialisation it makes it feasible to attempt error recovery.
* Explicit serialisation and deserialisation stages make it possible to construct persistent state that is portable across environments and through development
mmap() is blindingly fast, but it is a tradeoff between performance, complexity and reliability - naive use across the board is unlikely to improve all three!
Even without mmap, if your program successfully writes an invalid data structure to disk, the next time you run and read it back, you will crash. That's orthogonal to the method used to write and read the program state.
Your point about serialisation occurring with a known state is not lost in an mmap approach. If you have transactional writes (as LMDB does), then writing to the mmap also only occurs with a known state.
In reality, e.g. in OpenLDAP, you wind up with a minimal serialisation step, but with zero deserialisation. E.g., if I want to store a record: typedef struct foo { int len; char *data; } foo;
The struct and the data of the char ptr may originally have been obtained from two separate mallocs and thus reside at very different locations in memory. My serializer would simply allocate a new block and store the character data contiguously with the end of the struct. If you mmap at a fixed address, then you can store the address of the character data directly in the struct's "data" field, and the next time you come around to use this record, no deserialization is needed. Explicit deserialization is truly just a waste.
If you only use mmap, there's a risk of some corner of the object graph getting subtly wrong owing to a bug in one version of the software, and never getting repaired.
Versioning of data structures is also a problem.
I'd leave this pattern for use cases for which copying of memory on load has a measurable impact on performance, and where required integrity of data isn't very high - ideally, the data should be able to be regenerated from primary sources.
One case I know of where this technique was used was the precompiled headers file for the Borland C++ compiler. There, it makes sense on both counts.
One obvious example would be filesystems. They're basically specialized databases which treat your whole disk as a single gigantic file, and of course there's a long, proud history of programs used to repair corruption in them due to bugs or other problematic events.
To me, that's an argument for not using formats with incremental formats if you can get away with rewriting the file each time, but once you have enough data to where you can't afford a total rewrite each time, does mmap make the problem any worse?
If you're updating files in place without a whole rewrite, then you are effectively writing a filesystem.
I prefer to use a different approach where possible: write a diff log, and occasionally do a rewrite of the whole file.
When you do have to treat the on-disk file as a filesystem, I'd explicitly treat it as such. Have a low-level structure that does not need modification for different versions, is portable and highly tested. Do application-level structure with normal read-modify-write, but layered on top of the virtual filesystem layer. Consider having two modes, one where the virtual filesystem just proxies the real filesystem, and instead of a "file", you have a directory structure on disk. Consider using an existing component for this (zip archive libraries are popular, but not particularly efficient for writing).
Mmap does make the problem worse. It encourages you to be lazy towards versioning and portability.
LMDB is at the heart of the ubiquitous LDAP ( OpenLDAP ) and is very well optimized ( look at his benchmarks ). Now they are optimized for reading, which is important.
I would imagine mmap-ing with large amount of write will result in unpredictable performance....
For more general use, a simple way to keep the amount of 'pending writes' down to reasonable levels is to have a thread walking your mmap()'d regions and enforcing an msync() on a regular basis. This also gives you guarantees such as "each part of mmap'd file is no more than N seconds out of date".
Of course, you'll potentially increase disk load overall, since a page dirtied three times, once every second, might only get written out once by the kernel (after 5s) but if you're syncing every 1s it'll get written three times.
But you should be able to take control.
"This library also supports distributed, durable, observable collections (Map, List, Set)" "It uses almost no heap, trivial GC impact, can be much larger than your physical memory size (only limited by the size of your disk) and can be shared between processes with better than 1/10th latency of using Sockets over loopback."
http://www.boost.org/doc/libs/1_55_0/doc/html/interprocess.h...
"Boost.Interprocess offers useful tools to construct C++ objects, including STL-like containers, in shared memory and memory mapped files"
The mmapped files are going to break every time you make the smallest change to your data structures. This strikes me as not terribly useful for rapid development.
On our game project we use a data tool that allows you to create tables of data out of columns with specific data types as well as references to other table rows.
When you export from the tool it creates data files but also generates the C++ struct that represents the schema. The loader simply mmaps the data in to memory. Variable length data like strings gets packed at the end of the table with the fixed length records containing relative pointers in to it.
The data on disk is stored in the exact format that is needed at run time.
References to rows in other tables turn in to pointer like objects so you get a perfectly natural syntax as if the data had been loaded in to hand written classes.
It's perfect for rapid development because designers can add new columns or tables of data and then just tell the programmer what they are called.
It also loads lightning fast and has excellent run time properties due to cache coherency.
I've worked on financial transaction processing systems using memory mapped files as their primary means of data storage. It is very effective.
https://code.google.com/p/codesearch/source/browse/index/mma...
This give string,set,map,list,deque,vector types that work in the shared space.
http://www.boost.org/doc/libs/1_55_0/doc/html/interprocess/a...
However, it is insanely fast when it works.
It's not that your program crashes, it's that your memory becomes corrupt and that causes the crash.
Restarting program is essentially redoing all the memory. If you persist your memory your program will crash on run.
https://plus.google.com/u/0/+KentonVarda/posts/NKUUzx2nEsN
If you're looking for an easy way to exploit mmap in your code, Cap'n Proto is a serialization format that works similarly to Protocol Buffers but is designed to work well with mmap():
(Disclosure: I am the author of Cap'n Proto... and also the former maintainer of protobufs.)
While true and exceptionally powerful TLBs are limited in size and, being essentially a hash table put in to your CPU by either Intel and AMD, are likely optimised for common-case mapping patterns. Of course, these penalties still pale in comparison to disk I/O... but kind of a downer if you're dreaming of something crazy like a linked list where every node lives in a separate random page.
Optimizations aside, beware this approach. ASLR is one of your two best friends (the other is DEP). When you purposely circumvent the protection it provides a security researcher somewhere will make you the topic of a very pointy blog post.
Fair point about I/O failures though.
You do not need to call mmap directly, because malloc() will do it if your allocation exceeds MMAP_THRESHOLD (usually on the order of a megabyte). So you get this optimization "for free"--the only important part is that you do not call malloc() for each tiny object.