Mio – Cross-platform header-only C++11 library for memory-mapped file IO
github.com
github.com
Wow, I did not expect this. I'm really touched. I wrote this as a small utility for my own consumption because I was unsatisfied with the existing selection at the time, so I'm both surprised and delighted to learn that people are finding it useful. Although to be completely frank, I think this library is way too small and insignificant to deserve a spot on HN's front page, but it definitely made my day. So thank you kind stranger who posted it!
I am sure this won't be the last top HN post about one of your projects.
Perfect is the enemy of good.
https://www.boost.org/doc/libs/1_68_0/libs/iostreams/doc/cla...
I've always wanted to try irt, but it can't handle unexpected process failure (i.e. a crashed process will leave the memory in an unknown state) which is something I always end up needing.
I've seen parts that weren't much better then someones lib on github, because essentially that's what boost is.
Creating a new memory mapping can be pretty expensive! On both Windows and Linux, it involves taking a process-wide reader-writer lock in exclusive mode (meaning you get to sit and wait behind page faults), doing a bunch of VMA tree manipulation work, doing various kinds of bookkeeping (hello, rmap!) and then, after you return to userspace, entering the kernel again in response to VM faults just to fill in a few pages by doing, inside the kernel, what amounts to a read (2) anyway!
Sure, if you use mmap, you get to look at the page cache pages directly instead of copying from them into some application-provided buffer, but most of the time, it's not worth the cost.
There are exceptions of course, but you should always default to conventional reads.
The application will get a signal (SIGSEGV/SIGBUS, can't remember), and no information about what the problem could possibly be. Most applications do not catch these signals and will instead just terminate.
Even if you do catch the signal there is a real challenge to know what caused the signal and to keep consistent book-keeping to be able to perform any sane action in response.
At a previous employer we started seeing this problem when scaling things in production which was no fun.
And it gets a SIGSEGV and segfaults.
https://www.facebook.com/notes/daniel-colascione/toward-shar...
I read the page you linked to, and just off the top of my head, trying to manage paging by catching SIGSEGV is how do you determine that it's in response to a real bug (say, dereferencing an undefined pointer)? In my opinion, by the time you get a SIGSEGV, you can't trust the program at all. While it might be nice to have a process handle page faults itself, I think a better API than signal() is required.
As I detailed in the doc and on libc-alpha, you really do need some kind of synchronous exception mechanism to match how real hardware behaves, and it would behoove libc authors to make this mechanism not suck instead of pretending that synchronous faults would just go away.
If you caught SIGSEGV on AIX 3.2.5 to manage a mapped NFS file, you deadlocked that filesystem. Great fun!
Thanks for the flashback!
I, for one, used memory mapping in the past to significantly speed up code that operates on gigantic data sets. I guess that your mileage does vary.
I ripped out the mmap codepath, fixing the mmap-related bugs, and for some benchmarks performance improved by a factor of 20.
Now, there are certainly places where mmap is awesome. E.g. if you can push the mmap semantics up to the application level, or you need the sharing semantics etc.
You seem very experienced, so I hope you don't mind a question. In my use case the files were as large as tens of gigabytes and I was creating read-only mappings of 256KB-1MB chunks in them, keeping the mmap handles around according to a cache policy and RAM usage limit. Do you think in this case using mmap could in theory introduce performance gains?
[edit: typo]
I appreciate the responses--learned something new today!
You can tell that you understand how modern OS memory management works when you realize that the OS "automatically read[ing] the pages...from disk" and "flush[ing them] to disk" on memory pressure is paging whether those pages are anonymous pages or mmaped file pages. :-)
[Edit: flushing dirty file-paged pages is analogous to swapping anonymous memory to the swapfile. Discarding clean file-backed pages is a bit like discarding anonymous pages that have been made unused through munmap, process death, etc.]
But to the GP's point: you don't need (except to conserve address space) to limit file mapping size. I think he really wants something like MADV_FREE. But it's complicated.
Another advantage of using application-managed caching is the ability to take advantage of things like huge pages (which can drastically reduce TLB miss rates), whereas with conventional mmap of conventional files, you're limited to regular 4kB (or whatever) small pages and associated management overhead. (There's no reason in principle filesystems can't use huge pages for page cache, but AFAIK, nobody does it yet.)
OTOH, kernel management of page cache allows for better integration of cache eviction with system memory pressure signals and allows for multiple users of a single file to share the memory mirroring the contents of that file.
> Do you think in this case using mmap could in theory introduce performance gains?
It depends. The right approach depends on a lot of factors, including workload and developer complexity budget. It's funny, really: the more experience you get, the less likely you are to say "$SOLUTION is the bestest evar!" and the more often you say "well, it really depends, so I can't give you an answer".
What really strikes me as needless is someone using mmap to read a 10kB ~/.myapplication.lol.ini file or something.
I <3 mmap
You might appreciate this toy Go package I hacked together: https://github.com/lukechampine/freeze
It uses mmap and mprotect to "freeze" Go objects; if you try to modify a frozen object, the program crashes.
Do you know of a way around that, aside from building your own patched kernel?
https://opengrok.libreoffice.org/xref/core/include/osl/file....
I think there are two issues. In C++, a header only library makes sense due to templates, which don't produce object code before specialization. In C, I have seen the same with entirely macro-based "libraries" that are really simulating templates in order to make abstract data structures.
The second case though, my theory is that there is a new generation of programmers who don't understand the traditional compile and link phases because they are used to other languages. It doesn't fit their expectation of how it should work so they bend the tool to their expectation instead of figuring out the old way. The very fact that people on here are saying a programming language should have a "standard package manager" demonstrates the cultural divide: this sounds a little nutty to an old time C person.
The disadvantage (and the reason that I'm personally not a fan) is that they bloat compile time. If it's a library you compile it once, and then it gets linked repeatedly, but a header only library is compiled every time any CPP file that includes it is compiled.
It may be that you have used a lot of template-based headers, which may compile nearly every time because they are literally creating new code every time a new combination of template parameters is given.
You are correct that I'm mostly talking about template heavy header files. There is a strong correlation between template based header files and header only libraries. The matter) latter generally means the former.
but are you going to mmap stuff in all your .cpps ? Most of the time when I use an external library, it does not get out of a single implementation file, so it being header only does not really make it worse
Once the infrastructure is sorted out, header-only has mainly disadvantages, like making it harder to separate interface from implementation.
There's a good talk on the topic: https://youtu.be/sBP17HQAQjk
I made a similar library for writing data into memory mapped files, also a self-contained header-only lib:
https://github.com/Morgan-Stanley/hobbes/blob/master/include...
This one also serializes a representation of the type structure of recorded data so that it can be safely concurrently mmapped and read either with the same code or with the generic PL/compiler that I've developed in this hobbes project.
Where unpredictable disk latency is a problem, we've got a similar header-only lib for logging into shared memory (then have another process to consume this shared memory ring buffer and dump it to disk for concurrent querying):
https://github.com/Morgan-Stanley/hobbes/blob/master/include...
This pipeline works well for having lightweight C++ processes feeding large volumes of data to generic query processes that we can run out of band to look at this data in various ways (with a Haskell-like query language).
We did hit a slight problem doing things this way that the straightforward representation of data (as in memory) for some cases just used too much space and too much time wasted in I/O. Basically for complex market data, where data structures aren't trivial and recording ~100GB/day makes it very awkward to keep around a few weeks of data for random querying.
So I also made this header-only lib to write data into these mmapped files with a simple compression method (I like to describe it as generalizing Curry-Howard to probabilities) that gives us much better throughput, much smaller files, faster query times, and still support concurrent constant time random access queries:
https://github.com/Morgan-Stanley/hobbes/blob/master/include...
It gives us compression ratios about the same as EOD gzip, but much faster and importantly works online and with these query use-cases we have with hobbes.
Anyway, maybe I should write up those details somewhere else, I just mean to say that this is a useful technique and you can push it very far and do many things with it in a very straightforward way.
We should not forget compilation times. A project I use depends on spdlog, a header-only C++ logging library. The thing adds almost two seconds per compilation unit to single threaded build times. And since logging is kinda used everywhere, the whole project takes forever to build (trice the build time it would have had without spdlog, I've measured).
What benefit is so great that it is worth killing compilation times?
The same idea can also be used for C++ headers (unless it's all template interfaces).
With such stb-style headers, you can even get better compile times, because you can include all header-libs into a single implementation source file, giving the same advantages as unity-builds (merging all sources into one file).
PS: origin of the name "stb-style": https://github.com/nothings/stb (although I guess the general idea existed before)
Personally I love being able to test something out with ease and then make it work with the build system if I like the results.
Earlier this year I wrote a Win32 program to solve an obscure issue for my employer.
My main concern was making it possible for programmers with limited knowledge of C++ to maintain the application. We are a C# shop and I want my coworkers to be able to maintain the application after I leave.
Here are some of the other things I did to make life easier for the maintenance programmer.
I documented how the program works, what APIs it calls, and added links to Pluralsight courses on ATL and COM. The program uses printf format strings so there is a link to a tutorial on them.
I added extensive logging with line numbers for every action that the program takes. You can reconstruct the control flow by looking at the log file. Yes, I tried doing this.
I picked a header only logging library so that it would be easier for a maintenance programmer to update it. All they have to do is update the submodule that contains the library.
This just isn't really something that can be cleanly abstracted to do what you want it to do, even if you can make the code "look" the same.
But again, it looks like a nice, clean, modern C++ library. Just IMHO misapplied.
It's not that uncommon. Even notepad uses it:
https://blogs.msdn.microsoft.com/oldnewthing/20180521-00/?p=...
Just a data point, but we use memory mapping (both with and without names) in our product that runs on Windows.
Unheared by you.
In a world consumed by Electron apps and Javascript, lots of people couldn't care any less about performant IPC, sharing data across processes, and multiprocessing in general.
Really? Most larger (Windows) codebases I know of use memory mapping.
This must be incorrect.
Jeffrey Richter's advanced windows, a very widespread win32 book, introduced this technique to a LOT of developers, just when windows was the hot stuff for programmers. I don't think I've seen a win32 code base >100 Kloc without mapping. Understanding and using it is one of those rites of passage that marks the switch from junior to medior win32 developer.