File handling in Unix: tips, traps and outright badness
rachelbythebay.com
rachelbythebay.com
A UUID, generated by the system's CSPRNG. Ideally through a syscall if the OS has one, so I don't need to burn a FD. You can't pre-fill all UUIDs without running out of space, and the CSPRNG prevents prediction, short of using gdb to attach to the process mid-name-generation, at which point you're on the other side of the airtight hatch.
I do wish linkat could do an atomic replace of an O_TMPFILE file descriptor. It'd avoid the whole random-name business, which is an ugly wart.
This reminds me of Jonathon Blow's rant on similar issues on XBox and PS2/3 where the API, instead of doing the right thing by default required the developer to spend days/weeks using it correctly and having their app rejected over and over and in this particular case there was zero good reason for it not to handle the issues itself rather than pass it all on to the developer.
There either needs to be a notion of partial writes (what the article describes) or a write that gets rolled back if interrupted (probably really hellish to implement ... Maybe impossible for something like a socket where the remote machine may have already seen some packets).
EINTR or SIGPIPE are famously criticized as poor mechanisms for this. However even without that, interrupting a blocked write would remain a tricky problem.
Writing to a socket is unlike writing to a local file, and cannot offer such guarantees.
Case study: Windows has a feature called "transactional NTFS" which I think may have grown out of the WinFS experiment. They duplicated a lot of filesystem APIs with "transacted" versions. I think the ntdll APIs were a little cleaner than the win32 wrappers - maybe at that layer you can set a transaction on a handle and issue normal fs calls? Anecdotally I heard it was flakey. I believe it's in a sort of deprecated state today.
As I understand it, the ext4 file system can have full-data journaling with the “data=journal” mount option (instead of the default, “data=ordered”, which is as you describe).
You get bagged on for saying this but it's just inexcusably shitty. And that's that.
Would it? I believe kill(pid, SIGTERM) interrupts it all right. Or do you want to do this without terminating the program? But then we are back to square one: the programs have to handle failed writes somehow. How? "Abort, Retry, Ignore"? And then again, most of this stuff happens deep inside libraries, and every library makes different choices in handling such errors, and when you combine them in a single program, you end up with a mess. As always.
EIO sounds like it would make sense but can only be returned for errors relating to modifying inodes, the EIO that makes sense for NFS to be gone would be occuring when you call close(2) on the fd, or fsync(2) which has very unpredictable behaviour when it returns errors anyways. Either way, NFS doesn't handle inodes, so an EIO cannot be returned.
But then again, POSIX merely states that while EIO shall be returned if "a physical I/O error has occurred", it "may also be returned under implementation-defined conditions". Heck, if you squint at the right angle, even ENXIO could be somewhat appropriate. I don't think that extending the meaning of EIO is worse than making unkillable zombie processes, to be fair.
Anyhow, what are the peripheral devices that support only synchronous access? I don't believe there are any such ones in the modern PCs, but maybe there are some in mobile devices? Old floppy drives?
And what happens if some bytes have made it to disk, network, or pipe? Don't you want to tell the caller how many, in case they want to recover somehow?
Congrats, you reinvented the status quo.
One improvement would be to separate the error indication from bytes written. Then you could have a write that fails, with nonzero output. Windows' API allows for this, but I don't know how common it is. It would probably be a source of bugs the same way these write(2) corner cases we are talking about are.
And yes, Windows API's WriteFile does actually return less lpNumberOfBytesWritten than was passed in nNumberOfBytesToWrite, but only when writing to a non-blocking pipe; the disk writes are either fully written or fully failed. Sockets use their own special API that generally follows the write(2) semantics (although there are some hilarious corner cases with internal buffers: e.g., you can make a 1 GB write to a write-ready non-blocking socket and the internal buffer will grow to accomodate it, but further write requests will return EAGAIN).
Don't invoke system calls directly. Use a library that handles all this stuff for you properly. Better yet, use a language that gives you a better interface than C does.
this illustrates the sad irony of unix's genius. All the people who think like Dave Cutler (or whoever) are eager to bring Cutlerian ideas to unix. But unix is where they are bringing these ideas because unix is what won, and unix won because it did not make Cutlerian design choices.
I'm not saying that an API call that did the right thing wouldn't be great for a programmer; just that there are a lot of (simple) things going on in the kernel and tradeoffs for a simple userspace security model and some of this comes with that territory. And neither am I saying that a Cutlerian model is theoretically unsound. But things play out practically and this is how things have played out, and observers along the way discussed it the whole time.
More to say for sure, but please forgive me. My comment is for people who understand what I'm talking about. Explaining to people who don't understand is a useful and noble thing, but this ground has been covered so many times over ("worse is better"), and it's very detailed and historical and somewhat subtle, and I have limited time atm.
Who's default are we talking about here? The developers desired state of default, or the users?
The default action of SIGPIPE is to kill the process, which is exactly what you need for pipelines to work correctly. Related article: https://news.ycombinator.com/item?id=22647539
It’s not even entirely accurate to say that that’s EXACTLY how pipelined programs should work; maybe they want to do some cleanup before they die. SIGPIPE was essentially created as a hack for naive programs that didn’t properly check the return value of write(). EPIPE should have been enough, it’s a perfectly fine solution.
I would argue that programs that aren't meant to be used in pipelines aren't well designed programs.
> It’s not even entirely accurate to say that that’s EXACTLY how pipelined programs should work; maybe they want to do some cleanup before they die.
If you need to do some cleanup before you die, then that's exactly what a signal handler is there for you to do. Nothing stops you from exiting immediately after a signal handler.
I challenge you to clean up a temporary directory, like doing "rm -fr $TMPDIR/tmp$SECRET/" except not by running another program, in an async-signal-safe manner inside a SIGPIPE handler which then exits.
Hint: You're not allowed to call system(), readdir(), malloc() or any of exec*().
SIGPIPE is like an exception. If you really need to, you can catch it (to do cleanup or similar, as a sibling comment notes), but otherwise the default behaviour is fine and likely what you want.
I've made use of a similar technique when reading input: in code where an EOF is a rare case, but reads occur in multiple places, you can simplify the main logic considerably if you create a function that contains the read and a longjmp() upon EOF to a specific handler which then uses state variables to decide what else needs to be done; it gets rid of all the "if EOF" duplication and puts all the special-case EOF-handling logic in one place.
That only maybe works if you write most of the application code yourself but it would be unrealistic in larger projects that have to do more than read from stdin and pop out data on the other end.
SIGPIPE should not exist, no matter what hacky ways you found to abuse it for your advantage in some edge cases.
errno is meaningful only after the function failed. Short write doesn't count as failure.
After a short write your only legitimate course of action is to advance the pointer and write again.
> The fwrite() function shall return the number of elements successfully written, which may be less than nitems if a write error is encountered
> if a write error occurs, the error indicator for the stream shall be set, and errno shall be set to indicate the error
Second time I googled something you mentioned in this thread, eager to know more. This guy filesystems, folks.
"Create a file adjacent to the target path using mktemp or similar.
Write your data to it
rename() it from the temporary name to the final name."
I believe this is prone to data loss when the system goes down and the cause is that POSIX guarantees on write reordering are loose. The rename can make it to disk before the data gets written.
I think it caused actual data loss with one of the fancier filesystems which actually pushed the limits of write reordering and exposed this (pretty common) scenario as buggy.
Edit: I found an article about this problem. "Ext4 and data loss": https://lwn.net/Articles/322823/
This is a syscall after all. First priority is to provide a very general interface and remain compatible, not developer ergonomics.
I've inspected the family of fopen() functions (fread, fwrite), and they don't handle EINTR. Do they delegate this task to the caller? Or that EINTR doesn't exist there? I mean, the error code itself is used, and by quickly skimming XNU sources I can see it's used in some NFS code, but can it happen when reading from or writing to local storage (hard disk)?
Edit: some minutes later, I can see that e.g. libstdcxx supports EINTR, so maybe it can happen on macOS as well (the xwrite() function):
https://opensource.apple.com/source/libstdcxx/libstdcxx-104....
Signal handlers: POSIX specifies that, after a process runs a signal handler which was triggered when it was in the middle of a system call, either the system call will automatically be resumed or it will fail with EINTR, depending on a per-signal-handler setting. If the signal handler was set with sigaction(), this setting is the SA_RESTART flag; otherwise it can be set by siginterrupt(), but the default seems to be implementation-defined(?). Anyway, it seems that both macOS and Linux enable restarting by default. (But well-written library code should take into account the possibility that something else in the process could have decided to disable it.)
Attaching a debugger (using ptrace): This interrupts system calls because the debugger expects to see a userland state. On Linux it seems to automatically restart the system call, but on macOS it produces EINTR. This behavior doesn't seem to be configurable on either OS.
On both OSes, when EINTR does occur in read(), fread() does not cover up for it by automatically retrying, regardless of whether it successfully read any bytes beforehand.
https://gist.github.com/comex/58ba394588478bb1d506a998850b94...
If I run the program and press Ctrl-C, it prints "interrupted" but then keeps going. However, if I add `siginterrupt(SIGINT, 1)` after the call to signal(), interrupting the read() produces EINTR as expected.
Interestingly, the glibc manual used to state that the default behavior depended on the presence of _BSD_SOURCE or _GNU_SOURCE, but in 2014 it was changed to state that the default is always to use EINTR:
https://sourceware.org/git/?p=glibc.git;a=blobdiff;f=manual/...
Since 99% of the disk I/O is cached, write either success or fails completely. It's beyond scope of the userspace application to know what and when data arrives on the physical disk. Same applies to Windows as well.
Using direct I/O is different story where this is the correct behaviour.
Tldr: it used to not work. It was fixed in Linux in 2004.
No idea about other unix-like OSs.
Ha. I have seen TCP sockets code which breaks in a "partial write" scenario, and even a "partial read" scenario.
Code worked fine on their own LAN. Broke for someone else, they couldn't see why...
They were very surprised when I explained that you don't always get one message per read(), because they had for years. Or maybe they hadn't, but assumed something else caused their occasional application glitches.
Files are hard - https://danluu.com/file-consistency/
This definitely felt like an API shortcoming, ended up writing a wrapper for write that did what I expected.