Are You Sure You Want to Use MMAP in Your Database Management System? [pdf]
cidrdb.org
cidrdb.org
Saying "many other DBMSs have tried — and failed" is a little weirdly put because above that they show a list of databases that use or used mmap and the number that are still using mmap (MonetDB, LevelDB, LMDB, SQLite, QuestDB, RavenDB, and WiredTiger) are greater than the number they list as having once used mmap and moved off it (Mongo, SingleStore, and InfluxDB). Maybe they just omitted some others that moved off it or ?
True they list a few more databases that considered mmap and decided not to implement it (TileDB, Scylla, VictoriaMetrics, etc.). And true they list RocksDB as a fork of LevelDB to avoid mmap.
My point being this paper seems to downplay the number of systems they introduce as still using mmap. And it didn't go too much into the potential benefits that, say, SQLite or LMDB sees keeping mmap an option other than the introduction when they mentioned perceived benefits. Or maybe I missed it.
Anyways it's not like OS will do some magic one a page needs to be flushed to disc. It all depends on how quickly and how nicely scheduling is done. Paper also goes into software aspects of using an mmap system. 3.2 and 3.4 seem quite relevant software problems.
Caches are pretty heavily customized for the database design because they control so much of the runtime behavior of the entire system. The implementations are different based on the software architecture (e.g. thread-per-core versus multi-threaded), storage model, target workload, scale-up versus scale-out, etc. The "custom buffer pool" isn't just a cache, it is also a high-performance concurrent I/O scheduler since it is responsible for cache replacement.
Even if you were only targeting a single software architecture and storage model, it would require a very elaborate C++ metaprogramming library to come close to generating an optimized-for-purpose cache implementation. Not worth the effort. The internals are pretty modular and easy to hack on even in sophisticated implementations. In practice, it is often simpler to take components from existing implementations and do some custom assembly and tweaking to match the design objective.
In either case, you still have to understand how and why the internals do what they do to know how to effect the desired code behavior. For people that have a lot of experience doing it, the process of writing yet another one from scratch is pretty mechanical.
Also, a lot of the interesting research stuff going on right now is on storage engines that look fairly different from a traditional RDBMS buffer pool centric storage engine.
mmap has one large advantage, which is that it allows you to share buffers with the operating system. This allows you to share buffers between different processes, and allows you to re-use caches between restarts of your process. This can be powerful in some situations, especially as an in-process database, and cannot be done without using mmap.
There are many problems with using mmap as your only I/O and buffer management system (as listed in the paper and the SQLite docs). One of the main problems from a system design perspective is that mmap does not enforce a boundary on when I/O occurs and what memory to manage. This makes it very hard to switch away from mmap towards a dedicated buffer pool, as this will require significant re-engineering of the entire system.
By contrast, adding optional support for mmap in a system with a dedicated buffer pool is straightforward.
In retrospect, I’d say it was something like 50% driven by Java memory mapping having insane issues on Windows, 20% Java memory mapping having insane issues on every platform, 25% me being a noob and 5% actual issues with memory mapping on Linux.
I think if could say “this DB is posix only” I would try memory mapping for the next DB I build
I'm very interested, what issues? I've occasionally considered using mmap for a few things, but never had justification to. I'm curious what issues I might have run into.
Did you use JNI/JNA to mmap or via RandomAccessFile?
I've implemented a modest database in both C and Java with memory mapping. It's hard to explain the feeling of programming both but in general: with java I found myself trying to add intelligence at the cost of complexity to make things run faster. In C I could get speed by removing overheads and doing less work.
Great! Please reach out to us after you learn (again) that this is a bad idea!
At startup, it would read everything into memory so everything was hot, ready to go. The "database" was intended to run on a dedicated node with no other applications running. It also kept a write ahead log for transaction recovery.
mmap hits limits far earlier due to the kernel evicting pages with only a single thread and much of the process needing global-ish locks.
Use zoned storage and raw NVMe, with a fallback to io_uring. You need a userspace page cache of some sort. Maybe randomly sample the stack of pages you traversed to get to the page you're currently looking at, and bump them in the LRU. Feel free to default to stream latency-insensitive table scan operations without even caching them to not pollute cache.
I recently got to talk to someone who had written a query engine on the JVM
One thing I thing they said is that:
> "The buffer manager should probably be rewritten to use the native OS page cache. When we wrote it originally this functionality wasn't easily available and so we used Java DMA (Direct Memory Access) instead."
I'm not familiar with what this means, would anyone be willing to explain more about OS page cache and how you'd implement something like that?Would really appreciate it
https://lwn.net/Articles/457667/
You typically have to go out of your way to _not_ use the page cache.
The alternative is to open your file with O_DIRECT, which makes your reads/writes always interact with the storage system and bypasses the page cache.
https://docs.jboss.org/author/display/TEIID/Memory%20Managem...
https://github.com/teiid/teiid/blob/master/engine/src/main/j...
https://github.com/teiid/teiid/blob/master/engine/src/main/j...
The people at Symas (the company behind OpenLDAP) implemented their own storage layer, and wrote a document about that: https://www.openldap.org/pub/hyc/mdb-paper.pdf
Obviously some kind of persistent store is needed. But in the modern world RAM is so cheap and so performant, and cross-site redundancy/failover techniques are so robust, and sharding paradigms so scalable, that... let's be honest, a deployed database is simply never going to need to restore from a powered-off/pickled representation. Ever.
The hard parts of data management are all in RAM now. So... sure, don't mmap() files. But... maybe consider not using a complicated file format at all. Files should be a snapshotted dump, or a log, or some kind of simple combination thereof. If you're worrying about B+ trees and flash latency, you're probably doing it wrong.
I mean, yes, putting a bunch of drives on a local machine is cheap. But that too is a circumstance where a mmap()-using DBMS is probably inappropriate.
This is really, really false. It would make my life a lot easier if it were true.
(The first job I had was at a 10-person startup in a coworking space, barely making enough revenue to break even, and still consuming a vast, vast amount of data, because the product involved constant streams of sensor data from a vast number of tiny cheap devices. People tend to forget that not every single business is a CRUD web/mobile app whose average user accounts for $20/month revenue against at most a few hundred HTTP requests and a couple megabytes of disk.)
Even that depends on a lot of factors. Just as an example, if it takes $10 to build and deploy a sensor, and it returns one number per second, then $1 of RAM can hold 2+ years of data before archiving it.
[0] https://static.googleusercontent.com/media/research.google.c...
32. Which I would expect to be overkill compared to the precision of your average sensor.
> Also, you need to identify which user it belongs to, as well as - in our case - which device and which sensor on that device.
Which is the same for big chunks of readings, so I'm assuming a system that's able to store that metadata once per several readings.
> How do you protect against bit flips?
I dunno, I was just saying the cost of memory. You can have bit flips on any kind of server no matter how you're doing it, and they might get persisted, so I'd say that's out of scope.
> average incidence of uncorrectable errors per DIMM
That average is highly skewed by broken DIMMs though. They found a mean of a few thousand correctable errors per DIMM, but they also found that 80% of DIMMs on one platform and >96% of DIMMs on the other platform had zero correctable errors.
> Your tolerance of errors may differ, but you'd probably want to replicate the data at least once. Disk seems a good choice, for the added benefit that you can actually survive a power cut without going bust
Sure, disk for backup sounds lovely and would be extremely cheap compared to the RAM. I wasn't advocating having only the RAM copy, just saying that depending on other factors it might be reasonable for a RAM copy to be the main analysis database even for sensor data.
The post that started this thread directly says you should be using persistent files as write-only dumps/logs.
A lot, bit a GIGANTIC margin of the needs for a DB are not close to this, at all.
Starting with sqlite, that is likely the most deployed, is impossible that this scenario could be used for it (and that accounting how small most dbs are).
It sounds fine, but...
What is the cost of restarts? You need to read through the log, apply it, etc.
That can take a LOT of time, especially on cloud disks
Article: "Don't use mmap because modern storage is soooo fast. It will BLOW YOUR MIND!!!"
You: "I'm in the cloud!"
.........oh. Nevemind.