Fork() can fail: this is important (2014)
rachelbythebay.com
rachelbythebay.com
For example, in Rust, you’d never get into this situation, because a decent fork ffi function would immediately convert -1 into a Result carrying an error, and properly check errno. Java and C++ would throw an exception, etc.
Thus preventing all sorts of bad behavior up the stack.
Some other call I don't respect enough would probably bite my ass, though.
The "canonical" example people use a lot is std::bad_alloc. Often, there is no point in catching it -- what cleanup or fallback work are you planning to do when you can't even allocate memory?
Of course, silently swallowing exceptions with an empty catch-block is terrible.
If you ignore or misuse error code, you carry on in a corrupted state.
Unfortunately people don't know about this possibility and write empty catch statements because "it's not going to happen anyway". Yeah right.
... and then you can just do `coredumpctl gdb` and start debugging right where things went wrong ?
Yay Rust, where things either fail (panic and quits the program) or returns a result that you can't use until you check for the error.
Fork (without exec) is a very sharp tool because it can violate ownership, doubly so in multi-threaded programs. So just because you get a Result from it doesn't mean all its pitfalls are handled.
Recently though I had to remove a bunch of these safeguards because I wanted to share handles and such across processes. So sometimes it’s nice to have that feature.
The typical way to report errors upwards is to create a pipe with CLOEXEC and then write the error code to it from the child if exec fails. This requires at least calling write().
You may be remembering vfork(), which is a bit of an obsolete API.
But yes, don't do this today, obviously.
So, yes, while working in Rust and using better APIs designed around these interfaces, we won't do the wrong thing, but there's still a lot of people that are going to end up getting cut by these old 70's interfaces. It would be nice if we didn't have to work with these, but many of us still do.
It doesn't predate creating a struct which tells you exactly whether thing went wrong and how.
… because you cannot fork() safely in Rust.
It is really wrong, but that is caused by the C-functions returning only one integer of data. And the bad C libraries. Null used to be -1 in some compilers. Mixing pointers as integers is a recipe for disaster, not compatible with some CPU-s, but it is C-standard.
In the windows32 API they improved it a bit, with functions that only returned an Error code. They required struct addresses in which results were placed. I think that the process structure is also a lot better than the fork() function.
Bonus points if the API you're using inconsistently mixes boolean, null-pointer, HRESULT, DWORD, and returning a status in a value you pass a pointer to (you did initialize that value, right?)
Why do we keep doing this to ourselves?
And do you know why you'd never get into his with rust? Because rust doesn't support fork.
"Separation" in this case implying "hard to misuse", which the POSIX fork() API certainly is not.
> And do you know why you'd never get into his with rust? Because rust doesn't support fork.
Yes, it does. https://docs.rs/nix/0.17.0/nix/unistd/fn.fork.html
Exceptions are a mess. Either you go the Java route of tagging everything as throwing every subclass of Exception under the sun (which encourages people to write empty catch blocks to silence a noisy compiler), or you go the C++ route where you are not totally clear an error can occur when writing the code or from glancing at it. (Combine with operator overloading for most confusing results.)
Having an error result that can perhaps be easily propagated to the caller is the best of both worlds, and I think is the thing that good C code tries to approximate in a more manual way.
Specifics vary by platform, but for FreeBSD i386, the result usually comes back in register EAX (but docs say sometimes another register), and the success or failure comes back as the carry flag. Of course, C never made access to the carry flag easy, so libc smooshes things together.
posix_spawn() {
...
pid_t pid = fork();
if (pid == 0) {
...
} else if (pid > 0) {
...
if (...) {
...
return ernno_from_child;
}
...
} else {
return errno;
}
...
}
I’m just pointing out that you’re still using errno, you’ve just hidden it behind a function. All of your syscalls are still getting their results stored in errno. It’s just fork() / exec() with some extra sauce, not a syscall in itself.I’m not trying to argue, I’m just pointing out that this is just a wrapper function on Linux anyway, and if you want to find an example of a different way to report errors from the kernel, it’s not posix_spawn().
Here is the relevant line in glibc:
https://github.com/bminor/glibc/blob/b5b7fb76e15c0db545aa11a...
Errors are returned as the function return value.
"The Linux implementation of posix_spawn{p} uses the clone syscall directly with CLONE_VM and CLONE_VFORK flags..."
Though you're right about errno, line 305.
int checked_fork() {
const int pid = fork();
if (pid < 0) {
throw std::runtime_error("fork failed");
}
return pid;
} pid_t pid = fork();
switch (pid) {
case -1: error
case 0: child
default: parent
}Raising awareness about bad APIs will also help people realize when they're about to create one themselves, which might cause them to switch designs.
To me, it seems better that there are fewer layers so that the architecture is clear, the elements of the OS and language are known, and the points where such errors are possible are clearly marked - as was done in this article submission.
And those problems have been this way for decades. The world you wish for doesn't exist. Not saying it shouldn't.
For someone who codes for a living and would be inclined to not check `fork` for an error code, would you mind sharing a bit about why you use that approach?
Usually when I write code, I assume it's my job to ensure the code works as designed in all reasonable circumstances. That requires understanding the failure modes of any external code I'm calling, at least as far as if/how it could fail.
IIUC, your approach is more like what I'd call prototyping. I.e., you accept that the code might be wrong, but you prefer to deal with that if/when it comes up in testing, so that you can move faster up-front.
Is that how you reason about it?
My suspicion (and naive hope) is that most people would not ignore that case if they knew it was dangerous. But I would expect them to ignore it if they didn't know it existed. Hence this article, I guess, which serves to remind everyone that it can in fact happen and is quite dangerous (due to a potential later interaction with kill).
I don't - the default behavior of simply ignoring I/O failures if, say, stdout's pipe was broken is usually what I want. In fact, I've had bugs in exception throwing languages where such stdout write failures threw, and I failed to explicitly catch and ignore them.
I also may skip error checking malloc. A null pointer exception / sigsegv / access violation "must" be fine if malloc is failing - even if I handle it nonfatally in our code, some of our closed source middleware doesn't, and neither do some system libraries. At best I can make slightly better fatal error messages for a subset of the resulting failures. If I'm trying to build super reliable software, I need to avoid exhausting/fragmenting memory badly enough for malloc to fail in the first place.
I have a decent chance of checking fork() for failure as I'm on the more paranoid end of error checking. I've seen enough weird junk like SetCurrentDirectory on a real directory failing due to NTFS filesystem corruption leading to an infinite loop - that I assume all documented error conditions will eventually occur somehow, as well as some undocumented ones.
But I've never seen it fail, and I'm probably just going to make it a fatal error.
I rarely use `malloc`, but I typically do check its return code. That's because in my experience, hunting down the root cause of segfaults can be a hassle, so I prefer to write code that will fail fast on failed memory allocations.
> That's because in my experience, hunting down the root cause of segfaults can be a hassle, so I prefer to write code that will fail fast on failed memory allocations.
This is a great reason to add error checking.
I've spent enough time hunting down harder dangling pointers and memory corruption bugs that alloc-returned-nullptr segfaults are trivial/second nature for me to figure out 99% of the time.
There is that remaining 1% of cases, for the already rare scenario where malloc has failed. Quite niche... but sometimes motivation enough for me to go through with adding error checking anyways. Especially if I actually encounter it.
If a failure would lead to data corruption, then yes. That this is important is made clear with gcc+glibc's _FORTIFY_SOURCE macro, which (e.g.) forces you to check the return value of sprintf() among others.
Of course, it's hard to imagine a case where printf() is part of a data pipeline that could be corrupted. So, nominally the answer is a simple `no'.
But it's a bad question, since fork() and printf() failures have different consequences. Also, boilerplate code for printf() never shows using the return value, whereas boilerplate for fork() always shows that. Your malloc() example is much better.
If malloc() failure leads to corruption, you'd better check the return value. Otherwise, sure, ignore it and let your prog just die is fine! Until it isn't ...
It's worth noting that "printf" specifically does not appear to be affected by _FORTIFY_SOURCE if the documentation I'm reading is to be believed.
> But it's a bad question, since fork() and printf() failures have different consequences.
Pointing out there are differences - via visceral example that likely hits close to home in your own practice as a programmer - is the entire point of the question, so I'll call it a good one ;)
The implicit point I'm making is that it's more complicated than "all syscalls must have error handling!" - which raises the question of, well, what does drive the choice anyways?. Ideally, perhaps a carefully reasoned consideration of the differing consequences you've pointed out. Or perhaps more likely, an adherence to examples and inertia. Or perhaps even more likely, a gut call based on things like "can I even think of when this might fail?", which one might answer "no" to - for having not yet developed a sufficiently mischevious imagination.
Regardless - having established the possible existence of nuance - this then lets us properly examine if our documentation does indeed properly prepare us to make these decisions, and how useful posts like the article might be for filling in where our documentation may have failed us.
> Also, boilerplate code for printf() never shows using the return value, whereas boilerplate for fork() always shows that.
Boilerplate for fork() does not always show that. As a concrete counterexample, the first google hit for "fork example" lacks error handling in any of their examples, although they at least mention the possibility of negative return values. I'll avoid linking it in an attempt to avoid improving it's undeservedly high search ranking, but if you're performing the search yourself, it's on the "geeks for geeks" website.
Maybe most sources of proper official documentation get it right these days, but I'm hesitant to assume even that.
> Your malloc() example is much better.
malloc vs fork failures also have different consequences. malloc failing is pretty unlikely to result in killing other processes (like your watchdog process!) for example. So perhaps it's just as bad of an example ;)
Disagree - the unfortunate interaction with kill is not clearly documented.
Do you implement error checking when calling printf, unlike every C codebase I've ever encountered which uses it? If not, you've implicitly acknowledged some error cases just aren't worth handling, or useful to handle. The question is then - when is it important?
PSAs like this one make it clear where the documentation may not have: Error handling fork() is important, and unlike error handling malloc where free(nullptr)ing later is a safe noop, kill(-1)ing later is an unsafe hazard that must be avoided. Additionally, it's frequently the case that the documentation is poor and would not help you even when you do bother to read it. Here's me previously ranting that the vast majority of documentation about atoi fail to clearly and adequately call out that atoi("a") is undefined behavior and citing my sources: https://news.ycombinator.com/item?id=14861447
> so how could someone fall under the impression that this wasn't the case?
Continuing past my atoi example...
Maybe they looked at alternative, poorer documentation. Maybe they looked at poor example code that didn't bother with error handling. Maybe they looked at decent documentation that failed to adequately stress the importance of error checking (EDIT: I'd argue this includes your wikipedia example). Hell, maybe they looked at great documentation - about a specific platform's implementation of fork, which perhaps makes fork() failing fatal to the calling process and thus "infalliable". Maybe they looked at the documentation for their favorite language's wrapper of fork, which throws an exception instead.
Maybe they didn't look at the documentation at all.
Maybe they learned of fork through word of mouth when the internet was down on a system without manpages. "Can fork ever fail?" "Hmm... I've never seen it fail." "Good enough for me!"
Perhaps this lack of knowledge can only come about by foolishness - but human nature and statistics mean at least one of your generally smart coworkers has probably fallen prey to such foolishness.
There's a reason that Unix sends EPIPE to a process when writing to a broken pipe, and why the default handler for SIGPIPE is to terminate the process. Interestingly, the Rust runtime blocks SIGPIPE, which was a naive and dumb thing to do [1], but which is now impossible to undo.
Similarly, in C the fail flag for FILE objects is persistent to permit alternative error management strategies. This is even carried over into Go, AFAIU. Basically, it's okay to leave a series of I/O statements unchecked so long as you check at the end of the block or transaction, or at the very least at close time.
The printf case is a bad example because the inconvenience of checking for failure on every call has already been accounted for.
I don't think I've ever seen C code that fails to check fork for an error condition, though I'm usually only ever reading my code and the code of widely used open source projects.
[1] Considering how much Rust touts the ease of FFI and integration with C and C++ projects.
As an author and consumer of cross platform software I can't rely on that - no SIGPIPE on windows - not to mention it's a pretty coarse hammer. Return value + errno is the way to go if I actually care about the exact behavior. One can wrap the call and panic/abort/terminate, but most people don't bother to do that, nor can I particularly blame them for not bothering (especially when I generally don't either.)
> I don't think I've ever seen C code that fails to check fork for an error condition, though I'm usually only ever reading my code and the code of widely used open source projects.
I'd be curious if git blame shows they had that error handling from day 1, or if it was generally added in response to heisenbugs. My money is on the latter, but perhaps you're generally exposed to better codebases than I am.
But shell programming and process pipelines aren't common on Windows. Rust shouldn't have changed the default behavior. SIGPIPE serves a critical role on Unix by ensuring that processes don't keep running when their consumers disappear.
> Return value + errno is the way to go if I actually care about the exact behavior. One can wrap the call and panic/abort/terminate, but most people don't bother to do that, nor can I particularly blame them for not bothering (especially when I generally don't either.)
Fortunately in Rust you're usually forced to handle the failure condition. The problem is with FFI to existing libraries that implicitly rely on default Unix semantics.
> I'd be curious if git blame shows they had that error handling from day 1, or if it was generally added in response to heisenbugs. My money is on the latter, but perhaps you're generally exposed to better codebases than I am.
Admittedly, I don't hang in the "I copied this code from Stack Exchange" or "IDE code completion is essential" crowds. When programming in C I always have the POSIX and C standards handy, even though I've memorized much of it already. I try to do the same for other languages, to the extent it's possible. Someone else linked to an earlier comment of theirs where they complained that the POSIX spec for fork was only the 8th hit on Google. Perhaps their problem was having a habit of Googling the usage for a core Unix API rather than using the freely available and eminently readable specs, which can be kept locally. I have well over USD $1000 of ISO specifications on my hard drive, and as far as specifications go the POSIX and C specifications stand out for their concision and clarity. This is true even when pitting them against the documentation for languages like Perl, Python, Go, or Rust.
Agreed, didn't mean to imply otherwise.
> ...
Agreed
> Perhaps their problem was having a habit of Googling the usage for a core Unix API rather than using the freely available and eminently readable specs, which can be kept locally.
To the degree that it is their problem, it's an understandable one. Not everyone has $1k to drop on text documents, and those standards may not cover the bulk - or even a particularly significant fraction - of the APIs one is dealing with on a day to day basis, making their value as an investment rather limited.
And two decades ago, I was gobsmacked to learn that C++ had a standard library at all, with it's own string type - std::string - that I could use where Borland's AnsiString was unavailable. Things have admittedly improved slightly since then, but you can't buy a spec without knowing it even exists!
And I'll admit I frequently google instead of refering to the standards I have bought. They're useful for resolving language lawyer disputes (although the occasional defect report limits even the standard's usefulness for that), but I won't call them particularly amazing references. Their main advantage is in being authoratative. Of course, implementations may not always perfectly match the standard they're supposed to match as well... when I care about the details, I'm more likely to dive into the implementation's source code rather than the spec.
Has this behavior with init always been this way? I could swear in the past that `kill -9 -1` used to kill init too (and thus cause a reboot), it was one of my favorite "fuck it and reboot" methods.
Other operating systems may vary, of course.
On success, the PID of the child process is returned in the parent,
and 0 is returned in the child. On failure, -1 is returned in the
parent, no child process is created, and errno is set appropriately.
http://man7.org/linux/man-pages/man2/fork.2.htmlI do agree that it was a deficiency in the way the course was taught, luckily that professor no longer teaches the course as it became obvious that he wasn't comfortable with Linux or C. Still, that's when it became apparent to me that if people will avoid using an environment they're uncomfortable with, they will. Some students made the grave sin of running their code in Visual Studio on Windows rather than using GCC, and it was accepted because the code worked. Those students graduated all the same, despite having no knowledge of man pages. It's a lot more common than you'd think (and more than I'd like).
More specificially, if you're going to use C on Unix/Linux, I think you have to know both.
> You can learn C syntax and how to compile it with pretty minimal Linux\Unix exposure.
Just because the "hello world" of the language is easy, does not mean the rest of the language should be just as fool-proof.
I read the HN comments from the first time this was posted, and the top comment told a story about some students experimenting with fork bombs. When the sysadmin found what they were doing, the text he gave them to learn better was, "A Commentary on the Sixth Edition Unix Operating System," colloquially called Lions Notes. [1][2] To me that makes the case clearly that that language and that set of OS tools are very closely entwined, and if present together should not be expected to exist in isolation.
It does sound like you had a deficient education regarding C. I think that means you needed a better education, however, and I think a lower level language built by and for programmers who know what they are doing should not be dumbed down to the level of badly taught undergrads.
Luckily, I was able to snag a copy of the K&R book early on and basically used that to teach myself the course materials. Unfortunately for the rest of my classmates, that professor's poor teaching made most of them dislike using C, so they suffered when we used it for Operating Systems later.
It may be outdated (I have no clue yet), but I'll check out the Lions Notes you linked. Thanks for sharing it!
kill -KILL -1
But I'm going to restrain myself :)As you have two processes after the fork call, you have to pass exec error results back to the parent process through something like a pipe.
> Neither of them fail often, but when they do, you can't just ignore it. You have to do something intelligent about it.
Nope. You can often do very dumb shit in response to it. Just like you don't have to do anything intelligent when malloc() fails, you don't have to do anything intelligent when fork() fails. The most common thing to do is to restart whatever was failing , starting with the smallest "thing" and going up in size (threads->pids->pgroups->containers->vms->bare metal, tasks->services->clusters, etc). Of course, you should probably have some kind of cumulative alert to tell someone to fix it, but often the restart is good enough to keep requests being processed. Such is the reality of imperfect production systems.
If you mean this to say you shouldn't check for errors, I disagree. Write errors that fail with an error code are much more common than those that do not. The latter in my experience is usually a particularly bad hardware problem.
For the most common write errors (including disk full, but also errors arising from hardware failures), I think you should make a best effort to check and flag it.
The larger your base of installed machines, the more you will be thankful for this logging. You don't want to be diagnosing rare failures blindly. A low failure rate multiplied by a large install base leads to that problem daily.
> Just like you don't have to do anything intelligent when malloc() fails
This is also false in my experience. I have seen stuff in the field where malloc(40) [yes, that small] fails, and I am still able to log where that happens, write it to disk, and have it be recovered and read by me, the developer. I've also seen malloc fail and a UI program survive long enough to report the error to the user and maybe even retry the operation successfully.
The reason for this has to do with how you react. In the first case: a log file write does not ordinarily result in memory allocation. [fprintf does, but there are other ways to log]. In the second case, firstly, mallocs that fail are often absurdly large, where a smaller malloc will succeed; and second, you will react to the malloc failure by unwinding your stack and this will have the side effect of many stack frames calling free(), which might give you enough space to keep going successfully.
Instead, you should measure the overall "golden signals" of application/service/system health, and if you find a problem with those, you can start tracing calls, which will show you the erroring system calls. Cross-referenced with other metrics, you find the root cause.
This is for when your application is part of a larger system, anyway. If you're writing a monolithic application where the entire system is your application and it can never go down, then checking and handling every error is more useful, but then your monolith starts to internally resemble a complex system and you're back to too much noise.
Keep in mind not every job will have you with total control of a machine as in a server use case. In many domains deploying code to someone else's device you won't know that you're running somewhere short of disk space, or whatever, until you see it in the error code.
Further, I suspect you are working in higher level languages where somebody has already wrapped the error codes in exceptions, and already logs unhandled exceptions to stderr. So what I would call very basic "error handling" is already done for you. When talking about C, if this more manual work is not done, ignoring error codes and continuing to keep on chugging is just stupid in a lot of cases, will cause you to silently lose customer data if you don't stop yourself on the error, etc., which is bad times.
>Do you kill(pid, signal)? Maybe you do kill(pid, 9).
>Do you know what happens when pid is -1? …
If you break these rules, programming is hard. Working with programmers who make unproven assumptions is also hard. Either don't hire them, or train that behavior.
Computers are good at following rules every time; humans are not. The systems should be designed so that the complementary strengths of each are used. This is one of the ideas behind types in programming languages: they provide checks that the computer (which is good at this task) performs every time you compile code. Put more simply, get rid of null from your programming language and you get rid of a whole class of errors.
(As for null values, modern type systems have no need for them. I guess you aren't familiar with them, because the alternative is not "shifting them to another arbitrary number". Checking out Rust / Haskell / Scala / Reason / Elm etc. might help.)