Sure, it might be more work than using mmap. But it's also more correct, forces you to handle edge cases, and much more amenable to platform-specific improvements a la kqueue/io_uring.
Aren't system calls such as madvise supposed to allow user space to let the kernel know precisely that information?
Madvise is discussed in the paper, and it notes specifically that:
* madvise is not precise
* madvise is... an advice, which the system is completely free to disregard
* madvise is error-prone, providing the wrong hint can have dire consequences
A user space buffer pool gives you precise and deterministic control of many of these behaviors.
In this case, there is the issue of who is the you on question
For the database in isolation, maybe not ideal
For a system whre db and app run on the same machine? The OS can make sure you are on friendly terms and not fighting
Same for trying to SSH to a bust server and the OS can balance things out
You are correct to an extent, but there are a few things yo noted.
* you can design your system so the access pattern that the OS is optimized for matches your needs
* you can use madvise() to give some useful hints
* the amount of complexity you don't have to deal with is staggering
Speaking as an OS developer, we're not going to try to optimize buffered I/O for a particular database. We'll be using becnhmarks like compilebench and postmark to optimize our I/O, and if your write patterns, or readahead patterns, or caching requirements, don't match those workloads, well.... sucks to be you.
I'll also point out that those big companies that actually pay the salarise of us file system developers (e.g., Oracle, Google, etc.) for the most part use Direct I/O for our performance critical workloads. If database companies that want to use mmap want to hire file system developers and contribute benchmarks and performance patches for ext4, xfs, etc., speaking as the ext4 maintainer, I'll welcome that, and we do have a weekly video conference where I'd love to have your engineers join to discuss your contributions. :-)
Another aspect to remember is that mmap being even possible for databases as the primary mechanism is quite new.
Go 15 years ago and you are in 32 bit land. That rule out mmap as your approach.
At this point, I might as well skip the OS and go direct IO.
As for differ OS behavior, I generally find that they all roughly optimize for the same thing.
I need best perf on Linux and Windows. Other systems I can get away with just being pretty good
The mmap/madvise approach works well for things like varnish cache, where you have a flat collection of similar and largely unrelated objects. It does not work well for databases where you have many different types of data, some of which are interrelated, and all want to be handled differently. If you can meet the performance needs for your product by doing what you're doing then great - that's a fantastic complexity saving for your business. But the claim that "you can design your system so the access pattern that the OS is optimized for matches your needs" is unfortunately not true. It might be good enough for what you need, but it's not optimal. That's why there's so many lines of code in other DB engines doing this the hard way.
If your database server is the only process in the system that is using significant memory, then sure, you might as well manage it yourself. But if there are multiple processes competing for memory, the kernel is better equipped to decide which processes' pages should be paged out or kept into memory.
"If you aren’t using mmap, on the other hand, you still need to handle of all those issues"
Which seems like a reasonable statement. Is it less work to make your own top-to-bottom buffer pool, and would that necessarily avoid similar issues? Or is it less work to use mmap(), but address the issues?
You get to take advantage of literally decades of experience
What is more, if you can match the profile of the optimization, you can benefit even more
Uou map the file once, then fault it in
- single-threaded calls to 'fallocate' will help avoiding sparse files and SIGBUS during memory write - over-allocating, caching memory addresses and minimizing OS calls - transactional safety can be implemented via shared memory model - hugetlb can minimize TLB shootdowns
I personally do not have any regrets using mmap because of all the benefits they provide
The downside is that writing an excellent buffer pool is not trivial, especially if you haven't done it before. There are many cross-cutting design concerns that have to be accounted for. In my experience, an excellent C++ implementation tends to be on the order of 2,000 lines of code -- someone has to write that. It also isn't simple code, the logic is relatively dense and subtle.
And let's not forget sqlite!
> There can only be a single writer at a time to an SQLite database.