Disk write buffering and its interactions with write flushes
utcc.utoronto.ca
utcc.utoronto.ca
On Windows, when I copied a file, disk writes started immediately. On older systems, like Win98, I had to tweak Total Commander's disk buffer to improve copy speed on the same drive. Total Commander even had separate settings for same disk vs. different disk copy buffer sizes.
When I switched to Linux I was immediately surprised that disk writes did not start until the memory was full, and then it would stop reading while flushing dirty data. This happens even if the copy is between different drives: reads stop, writes only, then reads again with no writes to the other disk, repeat. It basically halves the copy speed.
It even happens when I copy to network mounts: reads 20 GB of data into memory, then reading stops and tries to flush the data over the nfs. Nfs times out, transfer fails. I had to use nfs timouts of 1h just to be able to do a backup.
It drives me crazy. Is there any way to make it write immediately, or at least to put a memory limit on dirty data?
https://docs.kernel.org/admin-guide/sysctl/vm.html#dirty-bac...
The values are crazy high by default (on modern hardware anyway): 10% of memory for dirty_background_bytes and 20% for dirty_bytes. I wonder why no distro touches these.
Another set of people also complain Linux takes too long to safely unplug USB drives.
I've lost data as a side effect of a simple file transfer timing out.
There's no one-size-fits-all answer, which is why it's a tunable.
On my desktop with 32GB ram, I can even get audio to skip when ripping DVD's to disk. That's because practically the entire movie fits into ram before Linux decides to start the writeback process, and that writeback process will hog the disk for almost a minute. Or it used to, until I reduced the buffer size by a full order of magnitude.
This is just another sad example of buffer bloat: the inability to tune data buffers to the capacity of the underlying stream.
It's a real problem.
That's exactly what happens. The server ACKs data until it fills its write buffer, and then stalls unresponsive until the entire buffer is flushed to disk. If it takes longer to flush the buffer to disk than the client's timeout, it gives up.
I have personally watched this happen via wireshark where the server doesn't ACK for more than 10 minutes.
About 4-5 years ago, i was working on a project, and part of that was copying big amounts of data to a system via nfs. At 30 minutes exactly, nfs would croak, transfer fails.
I think this buffer fill and empty flow was fucking killing it. Its a shame i dont work there anymore, id definitely wanna try tweaking these settings and see if i could solve it
And it's the vm.dirty* settings to change to fix it as described here: https://lonesysadmin.net/2013/12/22/better-linux-disk-cachin...
You can confirm by watching /proc/meminfo and watching the Dirty and Writeback numbers.
Changing up the vm.dirty* settings can help as described here:
https://lonesysadmin.net/2013/12/22/better-linux-disk-cachin...
Yeah, I/O blocks drive me mad every time. They are more noticeable on Windows though. Perhaps that's because Windows doesn't RAM-cache enough and I do a lot of USB I/O (USB NIC, USB drives).
> Another set of people also complain Linux takes too long to safely unplug USB drives.
If only it had an APIs to see how much RAM of a specific device is RAM-cached rigth now and visualize the progress of flushing that cache... Unvisualized long I/O (incl. caching) operations, let alone those freezing the UI, indeed feel bad and are a UX bug.
One of the key reasons I prefer Linux over Windows is Linux is much more rare to freeze, no matter the workload.
This is a typo, should have been "If only it had an APIs to see how much of a specific device is RAM-cached right now".
I wish I could configure Windows the same way: whenever it can use RAM to avoid an extra disk write/read - it should.
I had 16GiB of RAM which meant quite large swaths of dirty pages would become buffered while the SSD sat idle until writeback began. This would cause high-FPS full-screen recordings in particular to become backlogged and start dropping frames / audio dropouts. Just generally broken behavior for a desktop recorder, especially for a deferred-encode mode that's supposed to be optimized for minimizing system-wide effects/overheads during the recording.
The simple solution I found was to proactively initiate writeback regularly via fdatasync() on the cache fd. [1] I haven't decided yet if more should be done to constrain its buffer cache effects though. The cache files will be read back during encoding in post, so if there's enough RAM it can be desirable to enable reading them back entirely from memory instead of having to hit the disk again... but it would also be nice to let the rest of the system's processes keep their stuff in the page cache. memcg can probably be used to find a balanced solution, but I haven't done any experiments yet. Have any of you handled similar scenarios? What did you do?
[0] https://github.com/recordmydesktop/recordmydesktop
[1] https://github.com/recordmydesktop/recordmydesktop/commit/42...
I think having the trigger be size based rather than timebased is the real problem. Or bounds on both...
I probably don't want to buffer writes for more than X seconds, or let the buffer grow beyond Y% of ram. At least for the time based limit, you'd really want to be able to say start writing to disk when the buffer has data over 10 seconds old, but still accept writes into a new buffer, only blocking writes when there's a buffer being written out and the current buffer is too old or too big.
Though in ZFS' case it's not really a regular write cache as such, as it's used to minimize updates to its on-disk copy-on-write structure.
It's not out of hand a terrible idea to avoid flushing data to disk and there's no free lunch here as any workload you optimize for will have a different workload that suffers. People try to come up with general heuristics that work in most situations on consumer machines, but there's no one size fits all for all HW + use-case combos. That's why hyperscalars tune the kernel beyond that / have kernel developers writing code to optimize for their use-case. It's telling that the performance analysis in the article is pretty hand-wavy without any clear demonstration of a concrete problem.
As for sync, I believe the author is mistaken. You can do fsync instead which is more efficient as it only creates a barrier for writeback of the file descriptor rather than a system-wide sync. And invoking fsync I believe is more common than sync. You should be able to have multiple parallel fsync happening concurrently for unrelated files that don't block on each other so much (ideally the kernel would prioritize those writebacks and interleave for fairness, but I doubt it does).
Disk also might be in standby mode, which saves motor spin up and head load cycles too.
The goal is optimizing performance as seen by users. Long waits for disk flushes because the system wasted write opportunities is not good performance by any measure.
Modern storage hierarchy has become so convoluted precisely because of many conflicting goals: performance, efficiency, durability, cost, etc.
[0] https://docs.kernel.org/admin-guide/sysctl/vm.html#dirty-exp...
[1] https://www.kernel.org/doc/html/v4.19/filesystems/ext4/ext4....
Having such behavior be the default is, to my limited understanding, a bug in the Linux kernel.
When I write lots of data to a rusty device and don't sync, the kernel gives back more quickly. From there, I can move on with my work : everything is "as-if" the data were actually copied (except power-cutoff protection, indeed)
As such, I see this behavior as a feature, it allows me to wait less and do my work quicker.
Unless windows, on which I must wait for the whole data to be written on the disk.
So, when (if?) you notice that the transfer has failed with a timeout, do you still consider your work was done quicker?
If you are thinking of something like a HTTP download, the transfer has to be done anyway (and I must wait for it). However, I do not need that data to be fully written to the (potentially busy) local device
It times out, 'cp' exits in error, and my file is not transferred. I can reproduce this at will. I think I should report it as a linux kernel bug, to be honest.
I don’t see any stats/graphs or benchmarks in the linked article.
This is probably much less of a concern today, as NVMe drives - beside having many orders of magnitude higher IOPS capacity - also have (at least on paper) much better hardware support for high I/O concurrency. It may still make sense, even today, if your hardware (or stack) limits IOPS.
https://lore.kernel.org/linux-mm/45odvhgymm7fxsgwpewoiiggaok...
https://lkml.iu.edu/hypermail/linux/kernel/1005.2/01845.html
I implemented something along the same lines but a bit less spicy here:
https://gitlab.com/nbdkit/nbdkit/-/commit/aa5a2183a6d16afd91...
This is especially true when the thing doing a bunch of buffered writes is in a VM. If the VMM is buffering writes to the host fs, you get the described effects in the host OS and the guest OS.
Original: If you are running the host OS you often have little or no say about what happens in the guest OS. Even if you control both it is likely not trivial to get the apps to use direct IO.
It could even be harmful to use direct IO because that would mean that writes would not stay in the guest buffer cache, forcing what would have been a cached read or minor fault in the guest into what it sees as a physical read.
The written blocks are not going to be shared, except maybe due to KSM. But KSM would do the same if that data was in the guest’s buffer cache if huge pages are not used.
Yes, and that won't change because the hardware with it's own buffers behaves the same way.
Yes, and there's not (generally) any ordering constraint, either. The last thing you wrote may be persisted, and not the first, etc.
1. open directory (dir_fd)
2. create an unnamed O_TMPFILE file (file_fd)
3. write to file_fd
4. fdatasync(file_fd)
5. linkat(dir_fd, "", dir_fd, "file name", AT_EMPTY_PATH)
5. fsync(dir_fd)
This should guarantee that either "file name" will have the old contents or the new contents and no transient version is observed. The "old" mechanism is similar in that you write out to a temporary sibling and rename (these days you'd use RENAME_EXCHANGE w/ renameat2 to guarantee the atomicity or get an error) with the differentiating difference being that the temporary data could be observed on the filesystem / left around if you have a machine reboot. 5. linkat(file_fd, "", dir_fd, "file name", AT_EMPTY_PATH)