Fast Commits for Ext4
lwn.net
lwn.net
> When we look at the fsync() and fdatasync() man pages, we see that those system calls only guarantee to write data linked to the given file descriptor. With ext4, as a side effect of the filesystem structure, all pending data and metadata for all file descriptors will be flushed instead. This creates a lot of I/O traffic that is unneeded to satisfy any given fsync() or fdatasync() call
Does it mean that, under ext4, a call to fsync is essentially the same as a call to sync(2)?
¹ http://shaver.off.net/diary/2008/05/25/fsyncers-and-curvebal...
It's not a bad idea to maintain a local buffer that gives you a certain amount of cushion. I recently helped a team resolve the exact problem you had, with a similar solution. But excessive, unnecessary use of non-pageable memory is one of the things which induce early I/O contention, causing these stalls to begin with. (Consider an overloaded or errant process generating and buffering a lot of logging noise, precisely because the overtaxed system is under heavy I/O contention.)
To reiterate: you want backpressure, which means that you want a process which is exhausting limited resources to slow down or block. And you want that to transitively slow down or block upstream requests. Too many developers don't understand this and insert hacks to solve their immediate problem (e.g. closing a ticket complaining about intermittent SLA latency misses) without appreciating the broader issues, which at the end of the day just contributes to these problems.
One of the alternatives people attempt is to insert a gazillion knobs to permit dedicated resource allocation. But now you just have two problems, the second being figuring out what the magic values should be--a never ending and often intractable problem. This rarely ends well except for highly specialized tasks--e.g. a dedicated DB administrator who spends all day attending to and tuning a database instance.
That said, in the old days you mounted /var (and if you were super fancy, /var/log) on different disks to minimize unrelated I/O contention.
That paragraph caused a whiplash of emotions while reading, "Cool!!!... what???? Ug."
> With ext4, as a side effect of the filesystem structure, *all pending data and metadata for all file descriptors* will be flushed instead.
Is the article mistaken?
It seems like a good idea for client filesystem to mount synchronously for data integrity. I'm not aware of i/o being a factor here.
Of course this doesn't prevent applications mishandling of file objects.
However in ext4 there is no need to call sync unless you disable batch processing of synchronous writes.
It's understood, there's no black magic. When there's not enough contiguous space to perform a metadata copy (it's a copy on write filesystem, no in-place overwrites!) then it returns an error, even if there's still some free but fragmented space or some free space in the data block groups. A rebalance when metadata blockgroup total size diverges significantly from actual use should solve this.
For small filesystems there also is the option to choose mixed block groups, then there won't be the situation where there's still free space in the data groups while the metadata groups are full.
The database itself contains journaling, so one might choose to run with data=writeback or even directly against the block device if they were concerned about performance.
- Filesystem journalling is making robust changes to the data structures describing directories, files, and where files live, in units of atomic filesystem operations. For example, the filesystem journal may record "CREATE FILE", which translates to "update directory entry 1234 in directory block 5678, then allocate and initialize extent descriptor 9999, then write an inode at array entry 74234"
- Database journalling is making robust changes to the data structures describing the actual file contents, in units of atomic logical application operations. For example, a DB journal may record "INSERT ROW", which translates to "update block 123 of this index file, and 234 of this data file", application-specific relationships like that cannot be captured by the filesystem on UNIX.
(Note: NTFS is transactional on Windows. It's entirely possible to correlate independent writes and make them atomic, so on Windows at least, in theory a DB could exist without a separate journal. I don't know if this is used in practice). Even if it were in use, it places severe limits on the kinds of concurrency optimizations a database system could otherwise perform, because all of that stuff moves behind the curtain of the OS interfaces.
You can create the file, preallocate space, fsync the inode and the directory to ensure that it will be visible after a crash and then begin using the allocated space as a journal. Then you only have to fdatasync or sync_file_range whatever part of your journal needs to be persisted and those syncs can now be unordered relative to the filesystem's metadata journal without risk of data loss.
So data=writeback can be used safely, but you have to be very very careful about getting the syscall sequence right. Most applications implemented with sufficient paranoia and so are better served by stricter ordering modes and auto_da_alloc.
My solution has always been to just add a RAID controller with a battery-backed write cache, and then to disable barriers on ext4 and switch to journal=writeback. The data gets written to the cache "instantly", meaning less risk of data corruption (either at DB or filesystem level) from crashes or power outages. Saves having to worry about any of this sort of stuff.
But that's because they want to stop the world and write the data. Or more specically write the data they wrote. The issue is that more data is getting written than POSIX deems required, and that's the perf hit.
Depending on your usecase, libeatmydata or battery backed RAID can help, but I'd feel a bit skeptical of 'we want to lie about fsync() to our DB' for many workloads.
(Sigh)... Another reason to move away from Fedora.
Would you mind sharing what other features you've found pushing you away from Fedora and why? For my part, I'll admit, BTRFS and ZRAM both work against hibernation out-of-the-box, which is my most major gripe.
Since the hibernation image is neither encrypted nor signed, it's a potential attack vector to subvert the UEFI Secure Boot mechanism. Because UEFI Secure Boot systems are common (Windows hardware certification requires it exist and be enabled by default), it was decided to, by default, save the space on disk for disk-based swap and use zram-based swap instead.
It is straight forward to create a disk-based partition during installation, and it is configured to support hibernation, while still remaining subject to the lockdown policy.
They were all aboard with Stratis, but here we are, btrfs moving back in.
Ugh. This is why we can't have nice things. I really don't want the kernel's filesystem performance to depend on the number of different UIDs writing to the filesystem. That is insanity!
Ted Ts'o is just wrong here: performance should take priority over preserving the behavior of applications that rely on non-contractual implementation details of the Linux kernel. fsync should sync only the indicated file, and that's that. We can add a mount option to let users opt into the older, safer behavior, but we shouldn't suffer for essentially an eternity because somewhere, someone might have written an application that depends on an ext4 implementation detail.