Windows Memory Mapped File IO
jeremyong.com
jeremyong.com
A long time ago, I implemented some C++ classes that would hide most of the additional work and also took care of allocating new objects inside the memory mapped file. See: https://www.iwriteiam.nl/D0205.html#13MMF (Note that this implementation actually makes use of a slightly different approach where pointers are relative with respect to the first position of the file. This implementation has the limitation that you can only open one such store as a memory mapped file. It was only later I realized that it was possible to do without the offset. I never came to rewriting all the code.)
> complex object oriented structures if you make use of 'relative' pointers
You just described the original MsOffice file formats (and given many people who ever tried to parse them a PTSD shock)Then the operator-> is just (this + offset).
I've used this approach successfully in many projects. If you're worried about differing stls(probably a reasonable worry here), then use any of the boost containers.
[1] https://www.boost.org/doc/libs/1_86_0/doc/html/container.htm...
[2] https://www.boost.org/doc/libs/1_86_0/doc/html/interprocess....
Also, you can link against ntdll.lib directly. Manually calling GetProcAddress for a few functions isn't a tragedy by any means, but in this case, why bother?
Great article nonetheless!
> Unfortunately, I don’t have great answers here, and view this as a legitimate use case that warrants a dedicated code path using other mechanisms at your disposal.
Right. I think it’s similar on linux. Mmap makes sense until the data size is too large and tlb misses start to add up.
The article is describing a separate problem where you can’t issue concurrent reads although that doesn’t feel true if you make use of madvise.
Also, if the vm area is large, pages might be thrashing, requiring the kernel to do loads of work swapping things in and out and changing the pagetable struct.
I doubt that that's any different than allocating memory assuming your working set would be the same. Largely it should be unless the mmap design makes it easier to facilitate a random access pattern in your data structure that you would think twice about if you actually had to issue the reads yourself.
The bigger issue is that madvise & friends risk cross table shootdowns since Linux might not have relevant support for DONT_NEED other than unmapping (vs scheduling an unmapping to happen lazily). However, that isn't necessarily the case with Windows.
Linux just takes that as a hint to unmap immediately and this results in a cross tlb shoot down. Cross tlb shoot downs are crazy expensive. Windows might have better support for doing it lazily and not paying the cost.
It looks like they are headed to a multi-threaded pread implementation now [3] and someone has created a patch to tweak the current mmap fallback that uses pread to perform better in the meantime [4].
[1] - https://www.libtorrent.org/upgrade_to_2.0-ref.html
[2] - https://github.com/arvidn/libtorrent/issues/6667
I'm my various travels in I/O landed in windows meory mapped files have never been worth the hassle for my (fairly typical) workloads.
Would love to know when this approach is actually "better" for some definition of better.