How does Linux handle writes?
cyberdemon.org
cyberdemon.org
That habit was;
# sync
# sync
# sync
# halt
When powering down ye olde unixe maschine. It was so far back I just vaguely remember I'd read that somewhere; possibly in a ye olde unixe manual, but I can't remember exactly. I think back then it made sense, due to older spinning rust technology and there might have been some possibility of the OS not flushing any waiting cache to disk before powering down.
Even though it's probably not needed today, I still do throw a sync command in when powering down, say, my NAS with spinning rust drives, and my PCs/laptop with SSDs.
There is one thing where I 100% perform a couple of sync commands, and that's right before unmounting a USB memory stick - I usually witness linux Very Quickly Copying a file from PC to USB stick. "Aha! No, you're just caching that operation you sly dog you!" - the sync command on a CLI tells me when linux has finally actually copied the file onto the stick.
Recently I saw a USB stick's activity LED stay on long after unmount. I'm not sure if it was writing or something else.
The little microcontroller that does the actual write to pages of flash might actually be busy moving things around and doing other housekeeping after the disk is unmounted but while it still has power.
I wrote the firmware for a flash based system, and used powered-yet-unmounted time to save a few statis to the uC's internal flash.
However, LEDs become rarer and rarer these days.
But writes are still getting added after the sync. Maybe I'm wrong but it seems like sync has to not try to sync those latecomers, because otherwise it might run forever, with new writes showing up each time it waits for current writes to sync.
Assuming the above is true, it makes sense to sync multiple times. You converge towards being fully synced. The first sync might take extra long (say 30s) because there could be tons of stuff unsynced. But the next sync should only have those 30s worth of writes to sync. Let's say that takes 5s. Now the third sync only has 5s worth of writes to sync, and that takes 1s...
This is all built on some assumptions about how sync works. But it feels like sync needs to be able to terminate & these assumptions are the only thing I can imagine that makes sure sync can ever finish, without never ending changes keeping it open.
that would have to be an awful lot of data, or a very slow drive, perhaps a floppy?
your other assumptions are sensible, to an extent. but way back when, we typically had a script that sent a message to all ttys saying "system going down in 30 seconds, please log off" - in fact ALL the other systems i used back then (Dec10, VAX, IBM 4381) did this before the sync (or whatever they did).
Good Kingstons did 5MB/sec, top of the line ones did 11MB/sec.
My example used more illustrative than real numbers. Still, 10s seems not totally uncommon, which is only 3x off.
Also I'm around 70% sure remounting system readonly does equivalent of sync & not allow more writes.
https://bsdimp.blogspot.com/2020/07/when-unix-learned-to-reb...
Still no confirm or deny on Linux.
I was at +2 and now 0 after this post rose to the top. Am I bitter? No...
That's the only way it makes sense. The caller controls when they start the sync, so they can make sure that all writes they care about have been issued by then.
But it's impossible to define a temporal ordering between writes and the time the sync completes. That's not a well-defined point in time. There's always a non-zero delay between the time the sync proper completes and the caller gets back control.
Just `sync` also works but it's asynchronous, so you don't really know how long to wait until the writes are flushed to disk (especially for USB sticks which can be really slow when writing).
I myself learned a ton from it, but as I understand, it used to be the bible.
Never encountered it. I presume it was available in the UK, but my Unix courses in college didn't use it, and after college I started off as a CAD operator, using Tektronix SVR4 workstations with 4109 terminals running Teknicad, which is where I became quite intimate with SVR4 as I was a bit of a computer nerd ;)
With the interwebs there became a massive need for many many servers, and an attendant need for many many sysadmins, and their training/experience would be purely administration, with more need to glean more information from guides.
In that same time period, use of computers shifted to more personal computers, so the demand for datacenters was from a lot of Windows9x boxes (that needed administration and didn't get any). And then the shift to WindowsNT and then AMD64, so who was "coming up" and how much they knew of systems programming kept declining.
Now we've started to adminster administration, with a lot of files in /etc that say "automatically generated, don't edit this"
Evi Nemeth's books were feeding the internet growth, not the real old time sysadmin's.
Foxley's book was published in 1985. By the middle 1990s there were tonnes of books being produced, none of which was really "the bible", unless you counted ones that had "bible" in the title, like the Waite Group's "Unix System V Bible". (-:
They later asked why I sync all the time, almost always in threes.
"Huh?" Looking through history, I see that I do. Apparently it's been ingrained into my subconscious, like a sort of stutter, and I type it sometimes out of habit when I'm deep in thought.
Old habits really do die hard.
Anymore, I just buy the sticks with light indicators. It's not done if it's not pulsing evenly...
Raymond Chen has a tangentially amusing story that I just found while trying to google the win95 screen: https://devblogs.microsoft.com/oldnewthing/20160419-00/?p=93...
And buffers for efficiency aren't new to computers. Consider returning a book to a library. You drop it off in the return box. Eventually, a worker moves it from returns to processing, where they will then put it on a shelving queue. Where it will eventually be carted around and reshelved. Odds are high that you have some form of buffering before it gets to the return box, even.
Disk semantics are subtly different from S3-style blob semantics in important ways. Crucially the immutability of blobs makes life _much_ easier for the blob storage than the ability to write into the middle of a disk file.
Given the latency problem, I think it would be interesting and useful for there to be devices which expose a "blob" API to a bunch of Flash on a PCIe bus, and for there to be a suitable OS API for this.
Flash itself is only weakly mutable! You can only write in one direction (usually 1->0) to a Flash cell. To go 0->1 you have to do a block erase on a whole bunch of bits, which takes much longer.
It isn't hard to start to realize that there are some things you can ask the controller to do, that they just can't do. They can pretend to for you, though.
As I understand it, host-managed SMR drives for enterprise require the host software to do writes in the correct sequence.
We're partly there. NVMe has standardized a key-value interface for SSDs, but as with any compatibility-breaking new storage interface it's pretty much only used by and available to large cloud computing companies that can afford to rewrite their entire storage stack to take advantage of it.
However, if you're one of today's lucky 10,000, welcome to the curse of knowledge. On that note, I can only imagine the war stories that PostgreSQL developers could tell about disk syncing.
> How slow? An SSD is 1,000 times slower than memory. A spinning hard drive is one MILLION times slower. Disks are multiple orders of magnitude slower than memory!
Latency? Throughput? I don't think SSDs are 1,000 times slower than RAM in either category, assuming fairly typical consumer machines, but I could be wrong.
Also, thank you for finding that imprecise language. An original draft had stated that I'm talking about I/O duration, but I lost that language. I will re-add it.
My numbers are for "main memory access", "SSD I/O" and "rotational disk I/O" from Gregg System Performance (chapter 2). In retrospect, with all respect to Gregg, that wording is itself pretty imprecise.
I think what really matters here is I/O duration rather than throughput. That is, the reason we use the cache so much is because a single I/O is really slow against disks...
Everything in this business is like a great big onion of caching and buffering layers. Disks have them too. This is why you often see consumer SSDs have very fast writes only as long as you don't write continuously. Continuous writes causes their internal buffers to fill up and the real write performance to be revealed.
There's also factors like write amplification and overprovisioning that enter into this set of equations. But if there's any takeaway from this, it's that I/O benchmarking is hell.
Like, let's say you are doing some writing using the C standard library...
* the library itself will buffer your writes, to reduce the number of syscalls
* I think your bio objects might get buffered, via plugging, to reduce the number of block requests
* your block requests will get buffered, so that they can be canceled or merged (I'm actually not sure if this happens in the "software staging queue" part of blk-mq, or prior...)
* as you mentioned, the disk controller may very well buffer the actual write operations
Well there's a reason it takes a click or key press to register about as long as it takes sending a packet halfway across the world (or a few times around the world, depending on what software you're running).
You could compare memcpy to sequential write speed or something, but it'd definitely not tell the full story.
I ended up creating a bunch of smaller files containing essentially assembly instructions, of the form "put X at offset Y", and grouping adjacent write instructions in ~100Mb chunks. This is just sequential writes, which the OS will gladly help buffer and the disk likes. Then I go over the files one by one and evaluate the instructions. Now since the "random" writes are concentrated to a relatively small area on disk, the buffering works well.
Despite writing 2.5x as much data (and reading the data an additional time), this makes the write operation take like an hour instead of several weeks.
> ...but this file was dozens of gigabytes. Turns out if you do this the naive way on a consumer disk, it will take weeks and it will wear out the disk incredibly quickly
This is what C's "fallocate(2)" with "FALLOC_FL_ZERO_RANGE" or "FALLOC_FL_PUNCH_HOLE | FALLOC_FL_KEEP_SIZE" flags are for, along with "pwrite()"normal RAM only has like 10x the bandwidth of a high end SSD, though
The man page for glibc open(2) used to say, "Most Linux filesystems don't actually implement the POSIX O_SYNC semantics, which require all metadata updates of a write to be on disk on returning to userspace", but now it notes that "before Linux 2.6.33" it wasn't implemented correctly
I remember a rather eye-opening discussion on either the postgresql-hackers or lkml about the semantics of O_SYNC not actually being synchronous. In addition to the implementation details, which depended on the version of the kernel and the underlying filesystem type, it was also noted that underlying disk hardware also had caches which did not necessarily behave synchronously.
Some places I can find that cover these details include
https://www.mail-archive.com/linux-kernel@vger.kernel.org/ms...
http://milek.blogspot.com/2010/12/linux-osync-and-write-barr...
They tell the OS that the write has synced even when it has not as it has confidence that it will.
Yes, you can use caching to hide this, in main memory and on the device (safely and unsafely, depending on the design and whether it is an enterprise drive with sufficient capacitors), but you can do the same with SRAM vs. DRAM and it doesn't make DRAM as fast as SRAM, or registers-vs-SRAM, likewise.
This might sound inefficient but in practice consecutive reads are so much more efficient than random reads that the difference between reading 1 byte and 4096 bytes is negligible when the data is stored consecutively.
Consequently, when designing on-disk data structures that are accessed randomly (like Btree pages in a database, for example), it usually doesn't make sense to make them smaller than 4 KB.
(It is possible to bypass the page table cache using O_DIRECT, at least for some filesystems, but that has its own size and alignment restrictions; typically you can only read 512 byte blocks.)
It's way faster to fill a cache line with 64bytes in one go than it is to make 64 1 byte requests to main memory.
> What happens if you unplug your computer before the operating system writes the data to disk? Well, you will lose the data. It’s as simple as that.
At work I'm writing software for industrial machines. It's essentially a Linux system in kiosk mode running on semi custom hardware. We actually do have a power button that simply shuts off power (much to my dismay), and we have data loss all the time. We just make sure that whenever there's actually important data that needs to be stored (very seldom) we make it blatantly obvious to the user that now is a very bad time to push the power button, and sync before those warnings are removed.
But we've had boxes that were rendered unbootable because someone decided to push the power button the moment an apt-get upgrade completed and a lot of updated files were still living in memory only.
Firstly, modern filesystems are designed to be robust to random power loss, so it's unlikely that you will corrupt your filesystem in a way that cannot be recovered from. It does mean that some unflushed data would not be written, but it will only affect files that were opened for writing at or shortly before your system lost power. Some filesystems (including ext4) can offer quite strong consistency guarantees if you configure them correctly, though usually stronger consistency means lower performance in some scenarios.
Secondly, the Linux regularly flushes dirty pages to disk anyway (15 seconds on my system, see `cat /proc/sys/vm/dirty_expire_centisecs` for yours) so you cannot lose more than a few seconds of data in the worst case. This is particularly important for the case of "oh, I forgot to unmount the USB stick before pulling it out of my laptop".
Thirdly, most applications guard against data loss by using some form of write-ahead logging and explicit syncing. For example, if you edit a file in vim and then exit with `wq!` it will call fsync() before exiting, which causes the contents of the file to be flushed to disk immediately. In theory it's possibly that you lose power in the middle of the fsync() call, but it's extremely unlikely. (Of course it would be just as likely that you lose power right before issuing the write-and-quit command.)
The sort of data that is not explicitly synced is where it's not really important to keep the latest information. Think about log files that are appended to continuously, or for example downloads in browsers. Do you really care if you have to restart an incomplete download after a sudden power loss? Probably not.
These foundational things are pretty frequently skipped or glossed over. There's a lot of stuff above and beyond that needs to be covered.
For example, in 1994, you probably didn't spend a whole lot of time talking about concurrency or parallelism. Most software didn't need to think about those things.
But beyond that, kids these days have a very different relationship with computers than we did. They very VERY rarely interact with the filesystem (in my day, that was everything). Instead, it's a slick UX with apps to tell you everything to do.
Back in my day, even fairly novice computer users were familiar with the idea of mounting/unmounting stuff because they'd pop in a 3 1/4 floppy write stuff there, and pop it out. Formatting a floppy was a fairly common occurrence (for some reason).
Things are just different. On the one hand information about this stuff is WAY more available than it ever was. On the other, there's a lot more information and things are constantly changing. You can see that in old games like starcraft. There, the authors had things like their own custom linked list implementation. You'd pretty much never see something like that in modern software. You'd get funny looks with people saying "Why aren't you using a library for that? There's `ac`"
Mind you, there was the same issue with NIS.
No, that's not quite right.
It happened not to block this time, but the write was not "non-blocking". It could have blocked.
It's misleading to tell people that there is always a page cache involved.
If you care about data integrity, you don’t trust the disk anyway, so flushing to the disk isn’t enough.
(On ZFS and bcachefs you can get that by disabling syncs completely; they never reorder writes in a visible manner.)