Durability: Linux File APIs
evanjones.ca
evanjones.ca
What part of the POSIX specs are you reading?
"The fsync() function shall request that all data for the open file descriptor named by fildes is to be transferred to the storage device associated with the file described by fildes. [...] The fsync() function is intended to force a physical write of data from the buffer cache, and to assure that after a system crash or other failure that all data up to the time of the fsync() call is recorded on the disk."
https://pubs.opengroup.org/onlinepubs/9699919799/functions/f...
I would guess that 'intended to' is there very specifically to ensure that this statement really means... nothing at all. It can fail but if it 'intended to' succeed that's enough.
If your system crashes due to a power surge that breaks your hard drive... guess what it isn't going to give you back the bytes you wanted.
By giving instructions to the device that the device understands and provides guarantees for...
By your logic there is no such thing as a guarantee in this world, period. A meteor could crash into Earth and wreck everything.
You're being snarky. That's not allowed in the rules here and it's unkind even if it wasn't.
> By your logic there is no such thing as a guarantee in this world, period.
So what does the guarantee really mean?
That the data will be on disk... as long as the crash isn't too hard? What do you usefully do with this guarantee? How is it better than no guarantee?
It does not mean "we're just writing this superfluous text to clock in extra hours or pass a word count threshold".
Where have I said I thought it was about writing time or a word count? You've imagined that.
I once wrote the Novell NetWare drivers for a SCSI host adapter vendor, and part of Novell’s certification testing was to run I/O heavy tests on a system that would randomly cut power to the drive or the computer and make sure that this did not corrupt things beyond what NetWare could fully recover from.
Hardware vendors like host adapter makers and drive makers wanted Novell certification, so made sure that was possible.
They might also provide settings that sacrificed that for more performance, but they would make sure you could get decent performance using the full durability settings.
A flash device that isn't designed carefully enough could lose multiple megabytes of semi-random sectors if it gets interrupted at the wrong time, as a consequence of giant erase blocks.
IMHO, from having spent a fair amount of time on this, durability is a "fluid" concept, especially on desktop hardware (Microsoft had the right idea when they built Win Update on TxF), so if the data is important enough to worry about it, you should architect the system in a way that the durability of a fildes or drive is irrelevant to the overall system.
(shades of MS Fnd in a Lbry by Hal Draper here)
P.S. You don't even need a crash to induce the condition you're describing. Just call unlink() from another thread and suddenly fsync() will no longer be able to (and shouldn't) guarantee reachability...
Of course POSIX doesn't "require that filesystems not get completely corrupted in the event of a system crash", in the same way that it doesn't require that your RAM isn't faulty or that the computer manages to turn on at all after the next reboot.
A software specification and its guarantees only make sense in the context of properly functioning hardware and surrounding software.
In this case, that translates to "hardware doesn't randomly lose data" and "file systems do not randomly corrupt themselves", because that is the fundamental job of those systems.
fsync() is about durability, telling the POSIX-implementing OS interface to instruct the underlying software (e.g. file system) and hardware to sync the described volatile data into persistent state. For a system running on a ramdisk, that naturally means a no-op; for a network mount it means a roundtrip; for a PC with a hard disk it means that the described data is now on disk, and that it is there, uncorrupted, after an ordinary power failure. If you have broken RAM, a disk that lies about flush commands, or synced only a part of the intended data (forgotten to fsync the dir), then the POSIX system has still done its job regarding durability, and fulfilled its guarantees.
It does. It specifies how the filesystem should behave in a particular (non corrupted) way. There is no exception that this requirement is somehow voided if the system wasn't cleanly shutdown.
> It is explicitly intended that a null implementation is permitted.
> It is explicitly intended that a null implementation is permitted. This could be valid in the case where the system cannot assure non-volatile storage under any circumstances or when the system is highly fault-tolerant and the functionality is not required. In the middle ground between these extremes, fsync() might or might not actually cause data to be written where it is safe from a power failure. The conformance document should identify at least that one configuration exists (and how to obtain that configuration) where this can be assured for at least some files that the user can select to use for critical data. It is not intended that an exhaustive list is required, but rather sufficient information is provided so that if critical data needs to be saved, the user can determine how the system is to be configured to allow the data to be written to non-volatile storage.
Again, the way the normative language you quoted above is written makes it abundantly clear (for a standard) that fsync doing anything is entirely optional.
> If _POSIX_SYNCHRONIZED_IO is not defined, the wording relies heavily on the conformance document to tell the user what can be expected from the system. [...]
If you check your Linux system, you will see _POSIX_SYNCHRONIZED_IO is defined:
printf "%s\n" "#include <unistd.h>" "#ifdef _POSIX_SYNCHRONIZED_IO" "#error _POSIX_SYNCHRONIZED_IO is defined" "#endif" | gcc -fsyntax-only -x c -
and hence the rest of that paragraph isn't talking about this situation.(I think there was at least one more article by danluu on this subject but this one I remembered by name, having discussed it several times with colleagues when considering using something (sqlite) less ad hoc for temporary data)
This is technically right interpretation of what POSIX says but totally ignores the fact that in such state the application will know that this happened and should handle that. Because in the first place the rest of the write request did not happen at all and it will not happen unless the application retries the write(2) with rest of the data. Most applications simply ignore this situation which is the reason that this situation is guaranteed to not happen on any even vaguely unix-derived OS (and exactly this is the number one reason for error messages of the "Error: Success" kind).
Suppose you have created a file in a directory that is not readable. How do you get a file descriptor for the directory to call fsync on? It's an error to try to open a directory for writing.
There is one way to open a directory without read permissions, which is with O_PATH, but the resulting fd can only be used "to indicate a location in the filesystem tree and to perform operations that act purely at the file descriptor level". And indeed if you try to use this fd to fsync() you get EBADF.
I guess this is just an edge case of an edge case that most people aren't concerned about. One way to work around this issue with a bare minimum of compromise to security would be a setuid program that does nothing but open its working directory then fsync it.
In that case, you always have a handle on the directory.
Opening the dir for fsyncing separately after having written the file is similar to doing close()+re-open()+fsync() on a file, which does not provide as strong guarantees as direct open()+fsync(): https://stackoverflow.com/questions/37288453/calling-fsync2-...