Investigating Linux phantom disk reads
questdb.io
questdb.io
With that said, after reading the first paragraph I immediately searched the article for "mmap" and had a good sense of where the rest of this was going. Put simply, it is just really hard to consider what the OS is going to do in all situations when using mmap. Based on my experience, I would guess that a ton of people reading this comment have hit issues that, I would argue, is due to using mmap. (Particularly looking at you prometheus).
All things told, this is a pretty innocuous incident of mmap causing problems, but I would encourage any aspiring DB engineers to read https://db.cs.cmu.edu/mmap-cidr2022 as it gives a great overview of the range of problems that can occur when using mmap
I think some would argue that mmap is "fine" for append only workloads (and is certainly more reasonable compared to a DB with arbitrary updates) but even here, lots of factors like metadata, scaling number of tables, etc will eventually bring you to hit some fundamental problems when using mmap.
The interesting opportunity in my mind, especially with improvements in async IO (both at FS level and in tools like rust), is to build higher level abstractions that bring the "simplicity" of mmap, but with more purpose-built semantics ideal for databases.
Maybe, but table 1 in Andy Pavlo's paper shows 7 of 10 databases surveyed do still use mmap.
Furthermore, that paper clearly demonstrating the issues with mmap came out in 2022 and most databases have been around longer than that.
That mmap isn't the future is maybe more certain than that mmap isn't common practice today (because it does sorta seem to be).
I have to remind myself regularly that the world is larger than it seems.
The other day someone on r/rust said "Surely everyone knows what the rust programming language is by now. Can we stop introducing rust every time we mention it in a paper?". But no, obviously everyone in your circles knows what rust is. But you don't know most developers. And I wouldn't be surprised if less than half of working developers have heard of rust, even if everyone you know knows about it.
I used to rant against using socket.io at every possible opportunity. The library has (had) crazy bugs in its reconnection code. In the right circumstances the library would violate ordering and delivery guarantees, or it would lie about messages being received when they hadn't been. But no matter how much I ranted about it, and no matter how many hundreds of issues there were on github, far more people used socket.io than the (much more reliable) alternatives because socket.io had a pretty website, good documentation and it was taught at coding bootcamps. I think the only reason its not as popular now is that you don't need it now that websockets are available everywhere.
My partner says she imagines asking questions of our families when she tries to imagine what the average person thinks. But our immediate families are still a really weird bubble - every single one of the adults has graduated from college. (And weirdly, over half of that group have also taught at college.) That's still a really biased set of people. Finding an unbiased set is wildly difficult.
That is true. So with the modern browsers now days latest Chrome/Firefox (both desktop and mobile) can support websocket seamlessly? I guess socket.io is kinda like the jquery of ws then? It will take some time to phase out.
Yes. And this has been true for nearly a decade. The jquery analogy is exactly right. These days socket.io is simply an over complicated wrapper around a standard browser feature (websockets). Just like jquery, it will probably hang around as long as we’re alive through sheer stupid inertia.
There are corner cases where it's great, like when you have a file that you know is 90% in-core already and you don't care about errors. But overall, read() and write() are simpler, faster, more reliable primitives.
But for saving data from process-a I would still use write() using MAXPHYS sized blocks: I'm not sure mmap would use optimal write block size, and with write it is easier detect errors (like ENOSPC).
Another good use case for mmap is sharing read-only (or rarely updated) dataset among multiple processes.
But I wonder when there are decent reasons to let the OS handle file I/O through the demand paging system, since it's good at it.
This is generous. At least, it's not a great working assumption. If you know anything about your workload, it's often possible to do better by specializing slightly.
At first glance (doing a few text finds and a really quick read through the paper after clicking through that intro page with the poop emoji at the top and disregarding Recommended Music for this Paper: Dr. Dre – High Powered (featuring RBX)) that paper seems too short to adequately explore the topic.
On the other hand these same guys (the CMU Database Group) are an amazing resource and their youtube channel offers some great stuff [1] that would allow a curious person to explore the topic in greater depth if they dug into the papers of everyone who gave presentations at CMU.
clickbait: How many of the world's leading software engineers who addressed CMU students and professors about their successful database products rely on mmap? The answer may surprise you.
https://ravendb.net/articles/re-are-you-sure-you-want-to-use...
Agree on CMU being a great resource.
Why?
> that paper seems too short to adequately explore the topic.
The paper was published in CIDR (https://www.cidrdb.org). The paper submissions for this conference are meant to be short (typically 6-7 pages) to ensure that people can get their ideas out quickly.
When you just use a write call you provide a unit of arbitrary size, and if you've done your homework that size is a multiple of page size and the offset page-aligned. Then there's no need for the kernel to load anything in for the written pages; you're providing everything in the single call. Then you go down the O_DIRECT rabbithole every fast linux database has historically gone down.
On x86, and I think every architecture, when you write to a memory mapping that is not already backed by a writable page, the kernel is notified that user code is trying to write. And the kernel needs to fill in the contents of the page, which requires a read if the page isn’t already loaded.
It has to be this way! The write could be a read-modify-write instruction. Or it could be a plain store, but I’ve never heard of hardware with write-only memory with fine enough granularity to make this work.
The sole exception is if the page in question is all zeros and the kernel can know this without needing to read the file. This might sometimes be the case for an append-only database. I don’t know exactly what QuestDB does.
Also:
> As soon as you mmap a file, the kernel allocates page table entries (PTEs) for the virtual memory to reserve an address range for your file,
Nope. It just makes a record of the existence of the mapping. This is called a VMA in Linux. No PTEs are created unless you set MAP_POPULATE.
> but it doesn't read the file contents at this point. The actual data is read into the page when you access the allocated memory, i.e. start reading (LOAD instruction in x86) or writing (STORE instruction in x86) the memory.
What are these LOAD and STORE instructions in x86? There are architectures reasonably described as load-store architectures, and x86 isn’t one of them.
The RISC core of every x86 since PPro is a load/store machine.
In any case, my actual objection was to “LOAD instruction in x86”. It makes no sense.
This is very true. Perhaps there wasn't enough context to what the article is describing. The read problem started to occur on database that is subject to constant write workload. Data is flowing in all the time at variable rate. Typically blocks are "hot" and being filled in fully within seconds if not millis.
Zeroing the file is an option to try. QuestDB allocates disk with `posix_fallocate()`, which doesn't have the required flags. We would need to explore `fallocate()`. Thanks.
I would expect quite a bit better performance if you actually write a entire pages using pwrite(2) or the io_uring equivalent, though.
Messing with fallocate on a per-page basis is asking for trouble. It changes file metadata, and I expect that to hurt performance.
From a developer:
“As long as the OS & the HW doesn't crash, the data is safe thanks to the page cache”
This is so strange to me. A database that is non-durable by default. OK…
I still don't have an answer for him. It sounds just as strange to me too.
This reflection [1] came from the founders of RethinkDB, a competitor of MongoDB at the time:
"It turned out that correctness, simplicity of the interface, and consistency are the wrong metrics of goodness for most users. The majority of users wanted these three trade-offs instead:
- A use case. We set out to build a good database system, but users wanted a good way to do X (e.g. a good way to store JSON documents from hapi, a good way to store and analyze logs, a good way to create reports, etc.).
- Timely arrival. They wanted the product to actually exist when they needed it, not three years later.
- Palpable speed [...]. MongoDB mastered these workloads brilliantly, while we fought the losing battle of educating the market."
MongoDB narrowed things down for a specific use case, and became the best for that use case. This comes with trade-offs. MongoDB was probably not the best database for healthcare back in the days, but that is OK. It did the job very well for other use cases and industries. And over time, they fixed the issue around losing data and became more stable. Essentially, they made developers feel like superheroes, and over time improved their product, and eventually grabbed a massive market share.
[1] https://www.defmacro.org/2017/01/18/why-rethinkdb-failed.htm...
It used to be open source. It's not anymore
> The parent company is a listed company and worth $15BN, 3x more than Elastic to put some perspective.
That's purely a capitalistic argument and makes no difference to whether the product is any good. For example, there's plenty of "churches" that are richer than MongoDB Inc. and absolutely abhorrent and evil.
> This comes with trade-offs.
The only thing that required the trade-off of data loss was cheating in benchmarks in order to hoodwink naive potential users into using their dangerous product. MongoDB Inc. has always preferred to lie to their users. It is not a database company; it's a marketing company with a product they label as a database. And that's a smart way to make money, sure, because of vendor lock-in, but it's not a smart way to gain trust.
As engineers we bear responsibility for how our work impacts society. Mongodb may have made their investors a lot of money, but they did sloppy work and didn’t do right by their customers. That’s not a success in my book.
For some reason, many databases that overwrite data support disabling journaling and/or fsync. E.g. SQLite has "pragma journal = off". You can lose the entire database from an ill-timed crash, if one important page gets written but another doesn't. To their credit, it's not the default, and the documentation is explicit about this:
> If the application crashes in the middle of a transaction when the OFF journaling mode is set, then the database file will very likely go corrupt.
There are many cases, and ingest is one of them, where no being durable is fine. If you can either:
* Repeat the whole process on failure (which is assumed to be rare) * Recover from the failure without data corruption (distinct from data loss, mind)
In those cases, being 10x faster is very compelling.
Note that this is about ingest for bulk loads, while online transactions not being durable is a really bad idea.
For bulk load ingest, you can usually retry the whole operation. Not so for transactions.
[1] https://questdb.io/docs/reference/configuration/#cairo-engin...
"Imagine you work in a library where you store books on shelves. Your primary task is to take new books and put them on the shelves (write-only load). You don't expect to read the books often, so the number of times you need to open and read the books should be minimal.
One day, you notice that several books are being opened and read more often than expected, even though your main task is to put away new books. This is confusing and unexpected, so you start investigating why this is happening.
After some investigation, you find out that the library assistant (the operating system) is trying to be helpful by anticipating which books might be needed next and opening them ahead of time (readahead). This anticipation works well when there is plenty of shelf space (memory) available. However, when the library gets crowded (memory pressure), the assistant starts anticipating the wrong books, causing unnecessary book openings (phantom reads).
To resolve this issue, you tell the library assistant to stop anticipating which books to open (disabling readahead) when you're just putting away new books. This solves the problem and reduces the number of unnecessary book openings. The experience teaches you the importance of understanding how the library assistant works and shows that addressing unexpected issues can lead to improvements in the overall library system."
In some sense maybe this could be thought of as readahead in preparation for writing to those pages, which is undesirable in this case.
However, what confused be about this article was if the data files are append only, how is there a "next" block to read ahead to? I guess maybe the files are pre-allocated or the kernel is reading previous pages.
If the preallocation is done using fallocate or just writing zeros, then by default it's backed by blocks on disk, and readahead must hit the disk since there is data there. On the other hand, preallocating with fallocate using FALLOC_FL_ZERO_RANGE or (often) with ftruncate() will just update the logical file length, and even if readahead is triggered it won't actually hit the disk.
If the index block also got evicted from the page cache, then could reading into a file hole still trigger a fault? Or is the "holiness" of a page for a mapping stored in the page table?
It is quite possible the filesystem caches, e.g., the file extent tree (including holiness) separately from the backing inode/on-disk sectors for the tree.
If you are using mmap, that will express itself as a segmentation fault, which you really don't want.
You _need_ to allocate the file ahead of time, so you can properly behave there.
There used to be a sys-wide tunable in /sys to control how large an area readahead would extend to, but I'm not seeing it anymore on this 6.1 laptop. I think there's been some work changing stuff to be more clever in this area in recent years. It used to be interesting to make that value small vs. large and see how things like uncached journalctl (heavy mmap user) were affected in terms of performance vs. IO generated.
Does MADV_RANDOM disable both "readahead" and "readaround"?
Has this article being written using ChatGPT by any chance ?