O_DSYNC on Darwin is completely broken for the last 15 years
twitter.com
twitter.com
On Linux you can mount a filesystem and invalidate some metadata updates using options such as noatime. This is also depends on the mounted filesystem, such as FAT, which might not even support the updated attributes. Are these constraints honored in these environments with these flags? (sorry, too lazy to write code and check this)
There is no data on what is actually "broken" or why it might violate POSIX.1-2008. What metadata is actually important or impactful?
I am very, very skeptical of claims made by some database folks, due to past experience. Some Oracle products, and from what I understand, still do not understand that monotonically increasing means greater than or equal to (>=). The equal to is the problem, because calling an absolute value timestamp function to derive a "unique" value will always fail (with the same return value) when you get faster hardware and more cores. In a past life (~2000), I had to add code to bork an operating system to appease this very misguided notion. Last month (May 2022) I got (2nd hand) info that after a hardware update to faster hardware, duplicate unique row_ids were being created in an Oracle DB.
Without additional details, I find it hard to understand what the issue actually is. Is this a detectable temporal violation of order of operations for a live system, or some sort of fail situation where an abort / reboot occurs -- would love to understand how the external logging of this was carried out, because how was that synchronized / timestamped?
TLDR: https://twitter.com/jorandirkgreef/status/153231416960472678...
You can see that, for those block sizes, the throughput of O_DSYNC is the same as F_NOCACHE, e.g. 3121 MiB/s vs 52.05 MiB/s for libuv's durable fdatasync() which is really actually using fcntl(F_FULLFSYNC) under the hood here.
In other words, O_DSYNC is only going as far as the disk's own internal cache. It's giving nearly the same performance as a buffered write.
So the reason it is “completely broken” and not simply “broken” then, is that setting the O_DSYNC bit provides no durability at all. It's a no-op.
Also, to be fair, if you're familiar with the space, it's been known for some time that Apple's fsync() is not durable—it's not like this is coming out of left field. What's new here, is that O_SYNC/O_DSYNC have the same issue. That's why we're awarding the bounty.
This test seems to be based on how "understood" technology works and an expected outcome. Any technology that works better or differently than what you expect will fail.
Intel has already proposed MRAM -- aka persistent memory. If I am understanding correctly, your tests would claim that MRAM backed disks would fail your tests, they would perform way too fast.
These privatives seem to be super specialized and very dependent on what actually happens. The system call you reference, fsync() was introduced in BSD 4.2 -- 1983. There were no journaling filesystems in 1983. I think it is entirely reasonable to apply ~40 years of filesystem research to invalidate past system call practice and call it out as voodoo / blog fodder without real measurements on current filesystems and hardware and to also document if it fails or succeeds when the plug is pulled.
I acknowledge that doing this is hard work and time consuming. You are not going to get a 10 minute tweet from this, but really, isn't that the point of a robust filesystem?
That there's a $1,024 bounty being awarded should be reason enough to suggest more than “10 minutes” went into this. While the issue was first reported in a tweet, you should know that it came to us from an experienced database engineer who worked on both Google Spanner and FoundationDB.
The initial triage was to compare O_DSYNC with F_FULLFSYNC on the same device, because the results should be within the same order of magnitude, relative to the same device—we already knew that APFS doesn't make the same distinction between O_SYNC vs O_DSYNC that Linux does.
"Intel has already proposed MRAM -- aka persistent memory. If I am understanding correctly, your tests would claim that MRAM backed disks would fail your tests, they would perform way too fast."
Again, I think you're missing that O_DSYNC was compared relative to the same device, with Darwin's custom F_FULLFSYNC as baseline.
Finally, to assume good faith, you can also imagine that once we had the triage in hand, the issue would have been confirmed independently as “correct” by a third party in the best position to make that assessment.
They never got round to implementing unnamed posix semaphores: sem_init() will always give EINVAL.
They never got round to finishing poll(). It fails on devices. Because this includes /dev/null, it can make using poll() on stdin and stdout in general-purpose command-line tools dangerous. It used to fail on pipes too, but possibly this is now fixed?
It's actually the thing you use instead of O_SYNC because it's expected to deliver better synchronous performance, tantamount to using fdatasync() vs. fsync().
A broken fdatasync() (or O_DSYNC) is kind of a serious problem.
I think that for macOS so far, the focus has been on ordering guarantees, i.e. write barriers to reset the disk back to some atomic point in time after a crash.
However, for distributed databases like TigerBeetle, we also need durability guarantees, i.e. be able to know that the disk has the data down cold, that the disk won't “forget” data that the node has underwritten and ACK'ed externally.