fork() can fail
rachelbythebay.com
rachelbythebay.com
We didn't know, so we wrote the program and ran it.
This was on a PDP-11/45 running v6 or v7 Unix. The printing console (some DECWriter 133 something or other) started burping and spewing stuff about fork failing and other bad things, and a minute or two later one of the folks who had 'root' ran into the machine room with a panic-stricken look because the system had mostly just locked up.
"What were you DOING?" he asked / yelled.
"Uh, recursive forks, to see what would happen."
He grumbled. Only a late 70s hacker with a Unix-class beard can grumble like that, the classic Unix paternal geek attitude of "I'm happy you're using this and learning, but I wish you were smarter about things."
I think we had to hard-reset the system, and it came back with an inconsistent file system which he had to repair by hand with ncheck and icheck, because this was before the days of fsck and that's what real programmers did with slightly corrupted Unix file systems back then. Uphill both ways, in the snow, on a breakfast of gravel and no documentation.
Total downtime, maybe half an hour. We were told nicely not to do that again. I think I was handed one of the illicit copies of Lions Notes a few days later. "Read that," and that's how my introduction to the guts of operating systems began.
It's kind of weird that, while root has always had e.g. 5% reserved disk space on the rootfs for emergencies, one thing no Unix has ever done is enforce a 5% CPU reservation for root so administrators can "talk over" a cascading failure. I think this is possible just recently in Linux with CPU namespacing, but it's still not something any OS does by default.
However something that has been /possible/ for a while (but not in practice done) would be to elevate root process priority over other processes. Probably not done due to daemons needing to run as root (which is decreasing as they're able to drop privileges these days).
In theory this can give higher priority to a process, but if you cannot get into the run-queue at all (fork bomb), or the problem is in kernel space (e.g., I/O access, hang, or a kernel space loop), then it's not going to help you much.
Sure if you carefully made everything fork-bomb-resistant then a cpu quota would be a part of it. Container systems use fork bombs as basic test cases.
Totally limiting the CPU utilization of a group of processes requires more overhead than changing the scheduling priority since you must actively account for the CPU usage. CPU cgroups should do just that though and in most cases the overhead should be acceptable.
In your comment's parent, I don't think raw CPU utilization was the issue since kabdib mentioned fork and it was in response to a post about fork failures. The problems caused by a fork bomb are not limited to CPU utilization, see: https://en.wikipedia.org/wiki/Fork_bomb
In any case, there will likely always be some system call you can abuse to totally exhaust some resource of the kernel.
If this is true, I would expect there to exist one or more articles entitled "how I brought down my Heroku host-instance" or something along those lines. Anyone got some links? :)
In the old days, the Amiga operating system did use static absolute priorities for its multi-tasking. This meant that if a task with a priority of 1 wanted to use as much CPU as it wanted, then all tasks with a priority of 0 or below would be completely starved. This meant that you could boost a certain process (like, say, a CD writer) and get close to real-time behaviour. I was certainly writing coaster-free CDs on a much less powerful Amiga than a Linux box that constantly made coasters from buffer under-runs.
Linux, however, has virtual memory and "nice", which complicates matters. A process with a niceness of 19 will still take a small amount of CPU in the presence of another process with a niceness of -20. In the presence of a fork bomb, you may have a very large number of processes. If they all (by some miracle) have a niceness of 19, you still have very little CPU time left for a process with a normal or negative niceness. Infinity multiplied by a small number is still infinity. Real-time priorities are the only thing that will save you here.
You also have the problem of being able to actually change the processes' nicenesses. That requires CPU time, which you no longer have. You would be better off sending a kill signal. You also have a race condition - you obtain (from the OS) a list of processes that are running that you want to renice or kill. By the time you have iterated through each one renicing or killing them, new processes have appeared.
Oh, it's you. That story makes you even more awesome :)
I remember well the day one of the elder neckbeards handed me my own photocopy of the Lions books. It was enlightenment in pure form.
Sure, you could get in and kill a fork-bomb before it did anything bad. But two or three on the same machine? And when you've got a couple hundred machines? It was easier to just reboot and let the victims who were inconvenienced handle explaining to the guilty how what they did was bad.
Then there were the guys who would log into one machine in a lab, fork-bomb it, move to the next machine over and make a change to their program, fork-bomb that machine, and expect to iterate that process until they passed the assignment. Leaving a wake of pitifully flailing workstations behind. Ahh, good times.
No, real programmers were writing fsck.
:)
mkdir("/foo", 0700);
chdir("/foo");
recursively_delete_everything_in_current_directory();
Running as root, this usually worked fine: It would create a directory, move into it, and clean out any garbage left behind by a previous run before doing anything new.Running as non-root, the mkdir failed, the chdir failed, and it started eating my home directory.
This is such a stupid problem I run into a lot. Both alternatives (doing things with absolute paths vs doing things entirely with relative paths) seem to have a lot of downsides. Overall relative paths seems to be way better, but then you leave yourself open to problems which "rhyme" with the one OP was talking about.
I was citing rearranging the tree is a potential issue with absolute directories, not with relying on CWD - that certainly could have been clearer.
"Also, you're assuming AT_FDCWD won't be -1, which could be reasonable but isn't guaranteed afaik."
Interesting point regarding guarantees. It's not -1 on any existing OS that I can find (it seems to be -100 on Linux and FreeBSD, -3041965 on Solaris, -2 on AIX), and shouldn't be for precisely this reason, but something to bear in mind if you are working on something more obscure that nonetheless has these functions.
Of course, you shouldn't be relying on reasonable behavior from functions passed a bad FD in general. It's just nice to have the additional defense when that does get missed.
It doesn't turn into a mess if you handle it correctly from the start. The way we ususally handle this in large applications is to have one single class like 'ApplicationPaths' which internally figures out all paths needed. No other code uses paths directly, instead always uses paths relative to ApplicationPaths.AppConfigDir/ApplicationPaths.UserConfigDir/ApplicationPaths.ExecutableDir and so on.
In a sense, this is "still a CWD" - but the differences would be 1) you can maintain multiple at the same time, and 2) you could close it.
Ever heard of PATH_MAX and ENAMETOOLONG? You will if you're using absolute paths everywhere. Sigh
What is evil is a program changing its working directory. That's when it becomes an evil global variable, rather than a non-evil global constant.
I was wondering before if it would be interesting to have a filesystem with transactional locking of paths, though I'm sure the performance would take a hit. Would be kind of cool to be able to do filesystem operations without constantly opening yourself up to race conditions and requiring extremely defensive programming.
I agree that treating cwd as a global constant solves most (at least) of the issues, I'm just poking assumptions to see what ideas arise.
(let (dir "/foo")
(create-directory dir)
(with-current-directory dir
(delete-all-files-recursively)))
Factor recognized the value of dynamically scoped variables: http://concatenative.org/wiki/view/Factor/FAQ/What's%20Facto...A lot of code became much simpler because of that decision.
So if fork is behaving as documented, returning -1 is because of "fork failing".
At what point do API authors share the blame for a needlessly harsh punishment delivered upon a predictably common error?
I certainly prefer to work with systems produced by people tending to think it'd be their fault more often than not.
It's just as wrong to feed kill() -1 as it would be to feed it -48585 or "babdkd" (unless that is explicitly your intention). A simple sanity check of if [ "${pid} > "0" ]; is all that's necessary to protect against this behavior.
So, I would argue, the fault lays on the users, not the creator, for not understanding the API when all materials necessary to understand said API are freely available.
(with that said, I think it's safe to say, we've all been bitten by not fully understanding some function before)
This kind of mistake is godawful and should not be defended. (but it's correctly fixed through stronger typing, not through choosing -48585 as the code for killing everything).
Actually, only -1 is the code that "kills everything"
> There'd be far fewer bugs if everyone knew exactly how everything else works.
Perhaps I misinterpreted your meaning, because you seem to be advocating using programming and scripting languages without actually bothering to learn them. Of course this can, will and does lead to very bad effects.
The bottom line is, if you are going to use a function in your program/script -- please, read the docs and understand what is will return at the very least.
You're arguing that people should read the API before doing anything with it. Parent's point is that this class of error can be avoided by strong typing (eg, via algebraic datatypes), negating the chance that it would happen in the first place. Which, I think, is the right way to look at the problem. But certainly, if you do have to use a weakly-typed, unsafe language which does not provide this kind of guarantee, be sure to read the documentation twice.
Which doesn't mean you won't get bitten when it turns out that the person writing a library you rely didn't RTFM.
Is there a reason fork can't be changed to just crash the program on failure? Are situations where a program usefully does something other than crash on fork failure, more or less common than situations where a program fails in the way described in the article?
If you were forking in a high-level language such as Python, failing to fork would raise an exception which would possibly crash the program if left unhandled.
C does not have exceptions so return codes are used to indicate success or failure. This is true for nearly every function, not just fork. If you're not checking for errors in a C program, it's going to break in unexpected ways, and will possibly be vulnerable to exploitation.
fork has 3 possible return values: - 0 for the child process - a positive number for the parent process - a negative number if it failed.
If you look at the man page's "Return Value" section, it is extremely clear, see: http://linux.die.net/man/2/fork
"On success, the PID of the child process is returned in the parent, and 0 is returned in the child. On failure, -1 is returned in the parent, no child process is created, and errno is set appropriately."
YES. Situation: fork bomb, can't create new processes. How do you notice? Typically because you can't spawn new processes from your shell because fork is failing. With bash, there is a kill builtin - so you have a chance of cleaning things up (depending) if you have a shell open. If the failed fork kills the shell, then oops, you don't have a shell open.
$ mkdir /tmp/foo && cd /tmp/foo && touch bar.txtThis style is pretty prevalent when writing test cases in shell (or just shell scripts in general), e.g. when using something like sharness[1].
http://www.gnu.org/software/bash/manual/bashref.html#index-s...
$ mkdir -m 0700 char *dir = "/foo";
mkdir(dir, 0700);
if (chdir(dir) == 0)
delete_all_files();with-temp-buffer is another example: a macro that bridges the functions for "string manipulation" and the ones for "buffer manipulation", since you start writing stuff this way:
(defun replace-in-string (str from to)
(with-temp-buffer
(insert str)
(beginning-of-file)
;;; Here you can use all your normal text editing commands
(replace-regexp from to nil t)
(buffer-string)))
Lots of dirty manipulation, but from outside its a pure function, and doesn't change the editor state in any way after it runs. char *olddir = getcwd();
chdir(newdir);
try {
do_stuff();
} finally {
chdir(olddir);
}Also, getcwd has a size parameter these days, and of course you want to check if the getcwd actually worked.
Specifically, the code you wrote would behave the same if `dir` had lexical scope.
The scoping of the dir variable itself is irrelevant.
What the person meant when he wrote, "In those times I wish I could use the emacs lisp way," is, "In those times I wish I could use a lisp _macro_" -- particularly, one of those macros that makes a change, runs some code ("the body") then undoes the change.
Since all lisps have macros, the code above would work in any lisp -- not just Emacs Lisp. Among lisps, Emacs Lisp is famous for its dynamically scoped variables. Consequently, the specific reference to Emacs Lisp perpetuates the confusion that how variables are scoped has anything to do with what we have been talking about.
with-auto-compression-mode with-case-table
with-category-table with-coding-priority
with-current-buffer with-decoded-time-value
with-demoted-errors with-electric-help
with-help-window with-local-quit
with-no-warnings with-output-to-string
with-output-to-temp-buffer with-selected-frame
with-selected-window with-silent-modifications
with-syntax-table with-temp-buffer
with-temp-buffer-window with-temp-file
with-temp-message with-timeout
with-timeout-suspend with-timeout-unsuspend
with-wrapper-hook(Defensive Programming 101 really).
Also, don't rely on implicit state (the "current directory") for a destructive command, pass the dependency in:
recursively_delete_directory("/foo")
Reminds me of a time when I was working on a package system for an in-house linux os build and carelessly had my fakeroot directory set wrong in my configurations, which ended up treating my local root / directory as the root of the fakeroot, which is as bad as it sounds. Running the script over-wrote my entire /etc directory among other important and un-recoverable things...
Gladly, it was a development system so nothing crucial was lost. Needless to say, I am way more careful today in-part due to this mishap (and hours of setting up a new dev system!)
To avoid this, you have to either get a lock on that directory somehow (opening a transaction), or you have to try deleting the directory and then check the return code or catch any exceptions to tell whether the directory existed at the time of the attempted deletion.
This is why I found your "Defensive Programming 101 really" comment slightly arrogant (perhaps I misunderstood the intention). Writing correct programs is not easy and one should not mock others, because there is always something new to learn.
recursively_find_everything_in_current_directory() crashed due to a stack overflow when SetCurrentDirectory failed due to a single corrupt NTFS directory.
To write that safely, you'd test for the result of mkdir and cd to return success before deleting.
In a shell, you'd just go
mkdir /foo -m 0700 && cd /foo && ${deleteallfiles}"I was wrong about why malloc finally failed! @GodmarBack observes, in the comments, that x64 systems only have an address space of 48 bits, which comes out to about 131000 GB. So, on my machine at least, the malloc finally failed because of address space exhaustion."
I think that's your answer. A mutex error return probably indicates an application bug, such as double unlock. You should probably assert or abort on "can't happen" mutex errors.
Programmers are lazy. If they take the time to document an error return value, then you should probably heed their warnings. :)
#ifdef NDEBUG
# define VERIFY(x) ((x), 1)
#else
# define VERIFY(x) assert((x))
#endif
Then you can write VERIFY(pthread_mutex_unlock(&lock) == 0);
You don't need, however, to consider the possibility of your program continuing to run after pthread_mutex_unlock fails.Your code could just as well be written as
assert(pthread_mutex_unlock(&lock) == 0);
which of course has the added benefit of not inventing anything new, i.e. being standard and immediately understood by anyone who knows the language and its libraries reasonably well.Of course not unlocking the mutex in non-debug builds would be a problem.
Thanks, and sorry.
Be careful about "can't possibly fail".
1. stop everything
2. coredump (if applicable)
3. return -1 from main
This is a classic way to get your application exploited. Google did it (at least) twice in Android: once in ADB [1], and once in Zygote [2]. Both resulted in escalation.
Check your return values! All of them!
[1] http://thesnkchrmr.wordpress.com/2011/03/24/rageagainsttheca... [2] https://github.com/unrevoked/zysploit
Still, I agree with you 100%: check your syscall return values, especially security-critical syscalls like setuid!
[1] http://lwn.net/Articles/451985/ and http://git.kernel.org/cgit/linux/kernel/git/torvalds/linux.g...
Why are you treating the programmer like a machine? They're not a machine -- they're human. Regardless if they fully understand the API or not things should have have sane defaults for HUMAN FACTORS reasons.
Bugs will always exist. The fact that the Linux kernel has many bugs is just one example of a code base that has over a decade of work put into it by many people with high skill shows that bugs are inevitable.
The goal should be to assume people will do stupid things and make fatal behavior more explicit/difficult. Do we really need -1 for kill to do such behavior? How common is that anyway? It's a pretty destructive behavior, and probably should be removed from kill. The human factors approach would say if you really want that behavior then write a for loop to do over the list of pids, because it should never be within easy reach especially for such an uncommon scenario.
Apple's iOS API is similar. Try to insert a nil object into an array? Crash. Try to reload an item in a list that's past the known objects index? Crash. So instead of doing something sane like reloading the entire list, the user has a shit experience because off by one errors happen easily especially in front-end/model work [1 re: fb's persistent unread chat].
Not recognizing the human part of things leads to issues everywhere.. reminding me of this article on human factors in health care previously posted on HN [2].
Conclusion: design for humans and default to non-fatal situations.
[1] http://facebook.github.io/flux/ [2] http://www.newstatesman.com/2014/05/how-mistakes-can-save-li...
but remember that fork was not written in 2014. It was written forty-five years ago. I'm not saying it was a great API design decision back then, but I'm willing to bet that it seemed a lot less "wrong" at the time.
Where in the C manual is kill() described again?
It's a lot safer to fail fast and fail safe than to hobble along with possibly undefined state doing who knows what to the system and to the user's data.
Granted, the specific kill() API could have been designed better, but since it's very well established by decades of history, the burden of understanding it lies with the developer.
As you admitted, the API could have been designed better. That's my point. Things that break things for users, especially those that take down entire systems, should be difficult or require more awareness to do, such as by naming it killall() like suggested in this thread. The human factors approach recognizes that operators aka programmers aka humans will not just make errors, but predictability so. We can incorporate those predictions into our domain and change design patterns to match.
There are many known patterns of errors that programmers make: edge case errors such as off-by-one, null dereferencing, etc.
Using the return value from a function is a common pattern generally. Having a return value that mixes an actual result (PID) along with an error code (-1) is problematic in that they're both integers so there's no obvious handler for the error case, and thus this cascade can happen. Then you add the fact that fork() rarely fails and that compounds the issue in hiding it into obscurity.
Generally when a system is maxed out of resources all kinds of things break and fail in weird ways; it just so happened that this particular cascade was super bad and needlessly so because of API design choices that have never been addressed.
if (daemon && !test_mode) {
int pid = fork();
if (pid == -1) {
fatal_error("Failed to fork");
}
if (pid != 0) {
write_pid(pid_file, pid, !test_mode);
exit(0);
}
} else {
write_pid(pid_file, getpid(), !test_mode);
}
Phew! if (daemon && !test_mode) {
int pid;
switch (pid = fork()) {
case -1: /* Error */
fatal_error("Failed to fork");
case 0: /* In child */
break;
default: /* In parent */
write_pid(pid_file, pid, !test_mode);
exit(0);
}
} else {
write_pid(pid_file, getpid(), !test_mode);
}If you don't care about portability, it's an easier call to make.
(And if you really want to fall through, I could add a `fallthrough` keyword.)
edit: i know it's not totally analogous since removing the exit and not noticing there's no break is a lot more likely than just randomly removing a break, but the "let's prevent someone clumsy from screwing this code up in the future" argument always makes me laugh a little.
Is a break after an exit too far? It may well be, but it's less clearly too far than two breaks are.
if ((buf = malloc(buflen)) == NULL)
goto outofmemory;Just surprised by the inconsistency.
if (buf = malloc(buflen) == NULL)
goto outofmemory;
is actually equivalent to that code: if (buf = (malloc(buflen) == NULL))
goto outofmemory;
So, malloc gives you a pointer, which is compared to the null pointer, giving you either 0 or 1. And that is assigned to buf. Hopefully your compiler will warn you about the type error that spawns from such dark magic. if ((rc = pthread_mutex_lock(mtx)))
err("mutex_lock failed: %s", strerror(rc));
but decided to switch to a better-known function at the last minute and completely lost the point.I think this just came from being drilled in school on the importance of checking error return values. No matter how unlikely never assume that something can't happen. If it really can't you should at the very least assert on it.
[1] - https://github.com/jervisfm/W4118_HW1/blob/master/shell/shel...
http://c-faq.com/null/ptrtest.html
For practical purposes, the null pointer and the integer `0` are one and same.
For most practical purposes, which is why I said that it wasn't a terrible approximation. However, they can be distinguished on some architectures:
intptr_t x = 0;
void *p = 0;
x == *(intptr_t*)(char*)&p;
I can easily construct fantasy scenarios (involving more than a bit of Doing It Wrong) where this would be relevant. I'm not convinced it couldn't ever be relevant without Doing It Wrong, if one in fact needed to work on that kind of a system. assert(NULL == (void *)0);
is fine, but int x = 0;
assert(NULL == (void *)x);
is not.When I read this article I thought it was preaching to a choir. I'm actually quite surprised people programming C don't check for errors. That's the only way the functions can provide a feedback.
fork() returns pid_t type which is usually mapped to int32_t. For this type there's no equivalent of Perl's "undef", the -1 is standardized as an error in all system calls that return an integer.
As for the argument why not send -2 instead, well guess what? Other negative values also have a meaning. Negative values in kill send signal to a process group instead of a process.
It's not libc responsibility to predict all possible things the programmer can do. Also unlike perl, C doesn't have exceptions so it can't exactly quickly terminate on error showing what went wrong.
Imagine C throwing SIGSEGV every single time a function failed.
That's the problem there. Kill takes an argument that's either a process id or a magic number or a different magic number or.... Those should be different functions, and the special cases like "kill all processes", "kill all processes in this group",... should be some kind of enumeration type. But it's C, so...
Returning undef has the advantage of being something completely useless as a process ID for any other function call.
$ perl -we 'kill 9, undef'
Can't kill a non-numeric process ID at -e line 1.
[1] unless you "use warnings FATAL => 'uninitialized';"Processes requiring other permissions can/should be spawned by asking systemd/chron/init/launchd to launch them for you.
If fork() had thrown an exception for an unexpected failure then the user could not have accidentally ignored it in the same way.
I realize that this is not appropriate for a system call but it seems like a good example of why handling errors using exceptions is helpful sometimes.
Standard file handles are another thing you should not assume are there (though I'm not sure how to test for it programmatically).
We once had a user that, for whatever reason, tweaked their Unix installations to not pass an open stderr to processes - they just got stdin and stdout (that is, file handles 0 and 1, but not 2). If you wrote to stderr anywhere in your program, it wrote to whatever was open on handle 2, which was not a stderr that the OS passed in.
Yeah, that's a pretty insane thing to do, but somebody was doing it...
Note however, that strictly speaking stderr does not have to be 2. It can be any number, but it has to be whatever the include file says it is, so if you don't specify the stream as stderr but instead write to stream 2, that would potentially be a problem.
That said, if for some reason you have a program completely defying the C standard, you can test whether the streams are open (and they are explicitly defined as having to be open) using fcntl and testing for EBADF at the very beginning of the program.
Nope. On POSIX systems stderr is defined as 2:
http://pubs.opengroup.org/onlinepubs/9699919799/functions/st...
> The following symbolic values in <unistd.h> define the file descriptors
> that shall be associated with the C-language stdin, stdout, and stderr
> when the application is started:
>
> STDIN_FILENO
> Standard input value, stdin. Its value is 0.
> STDOUT_FILENO
> Standard output value, stdout. Its value is 1.
> STDERR_FILENO
> Standard error value, stderr. Its value is 2.stderr is defined by the C standard. POSIX is a standard followed by many of the systems that run C.
Strictly speaking, stderr does not have to be 2. It wouldn't comply with POSIX in that case, but it would still be C.
Section 7.6 explicitly says all three must be open when the program begins. There are equivalent sections in the formal standards, I cited K&R because it was on my shelf.
Please stop and bother to check your facts before continuing. You're very close to being accurate (POSIX does define open, but C defines stderr and some functions to print to it, the underlying mechanics are the choice of the implementing system.) But don't you think it would be nice to check that I actually am before writing yet another post simply asserting I'm wrong?
The sources you cite say nothing of file descriptors; they are all references to the standard FILE* interface in C. Those are opaque pointers and have nothing to do with 0, 1, or 2.
You may be confused because I abbreviated "standard error" as "stderr", yet I was never talking about the C standard global "FILE *stderr". That was sloppy of me.
POSIX implements this using file descriptors and specifies that stderr is 2.
I said: "strictly speaking, stderr does not have to be 2" which is true. A system is welcome to implement file descriptors and make stderr's something other than 2. It will be a blatant violation of POSIX, but complying with POSIX is optional. Systems that don't just aren't POSIX systems. Complying with the C standard? Not really optional.
fprintf(), but yeah. That is exactly what I've been saying, too.
> I said: "strictly speaking, stderr does not have to be 2" which is true.
Ok, but that's just like saying, "strictly speaking it's a valid C to write a bunch of zeros to a file and call it a jpeg."
It might be technically true (complying with the JPEG spec is optional for C programs, too), but it's in no way a reasonable thing to argue for.
As for the printf/fprintf flub, totally.
That's the thing. The C standard does not allow it. The C standard merely never mentions it, which is a different thing entirely. Hence my analogy to JPEGs (which the C standard never mentions either).
And I meant the argument itself is unreasonable. It makes no sense! I could just as well argue that the TCP/IP RFCs allow fd 2 to be something other than standard error. Or the HTTP 2.0 spec. Or the Ecmascript spec. Arguing that something is allowed by a spec that never mentions it and has absolutely nothing to do with it is not an argument.
https://www.kernel.org/doc/Documentation/vm/overcommit-accou...
NT does a much better job of separating these concepts than Unix-family operating systems do. Conceptually, setting aside a region of your process's address space and guaranteeing that the OS will be able to serve you a given number of pages are completely different operations. I wish more programs would use MAP_NORESERVE when they want the former without the latter. (I'm looking at you, Java.)
One day, perhaps when I am old and frail, we will achieve sanity and turn overcommit off by default. But we're a long way from being able to do that now.
One is if you set a lower process limit.
Another is if you are allocating lots of memory with alternating mprotect() permissions. On some systems (AIX for example) this uses up all of the memory for control structures WAY before hitting the address space limit (I've seen it fail after just a couple of GB).
rm -rf $PREFIX/usr/lib
in a Bash script being run as root. PREFIX was misspelled, and set -u was not in effect, so the misspelled variable silently expanded to nothing ... rm -rf /usr /lib/bumblebeeI identify myself as an atypical coder.
In a functional language this would look like a datatype (essentially a generalized enum), while an OO language you would use different subclasses of a common superclass.
In C you can return a tagged union and check the tag, but nothing forces you to do the check. A user of this API can just go ahead and assume the success branch of the union. Furthermore, this isn't idiomatic POSIX, so it is never done.
[edit: "subclass" -> "superclass"]
Tagged union is probably the best approach. It doesn't prevent skipping the check, but it at least makes "thing I am supposed to use" different than "thing I am supposed to check".
If you don't understand how to use sharp tools, you may hurt yourself and others. Documentation for fork() clearly explains why and when fork() returns -1. Those that find the man page lacking or elusive may get more out of an earnest study of W. Richard Stevens' book, Advanced Programming in the UNIX Environment. In any case, every system programmer should own a copy and understand its contents.
I'd argue every programmer. It's such a fundamental part of computers & operating systems that key concepts will come up again and again. Just the other day I wanted to learn about Docker/CoreOS/etcd only to realize that I have an embarrassingly lacking understanding of how UNIX works. I immediately went to the library to pick up this book and begin fixing a flaw of mine (even as a web developer).
But even if there was only UNIX... the entire point of a well designed system is to allow users of the system to reason about it on a high level, not a domain-expert level or even domain-intermediate level. As programmers we reason about code without worrying too much about gate layout on silicon. As non-system programmers we should likewise not need to worry about shoddy OS design.
It's still bad API design when naively handling an error case kills everything. Is there an inherent reason that the error value for a pid has to be the same as the "all pids" value? Unless there's a very compelling reason, it seems like very poor design, well documented or not.
This is the same sort of argument as strlcat vs strncat, and people can't agree on that one.
if ((pid = fork()) < 0) err_sys("fork error"); is idiomatic in Unix.
Once I discovered the existence of the BSD err()/errx()/warn() functions, though, the error handling in even my quick one-off programs became much better and more informative.
pid_t child = fork();
if (child < 0) {
err("fork");
} else if (child == 0) {
... in child
} else {
... in parent
}
is idiomatic, quick to write, and produces useful error messages when that "throwaway" program starts failing years later.See /etc/security/limits.conf and nproc and "fork bomb"
Aside from intentional fork bombs I've seen this done intentionally in the spirit of a OOMkiller to keep a machine alive for debugging / detection of problem. 100 "whatever" processes will kill this webserver making it impossible to log in and diagnose much less fix, so we'll limit to 50 processes in the OS.
I've also seen it in systems where people are too lazy to test if a process is running before forking another and the system doesn't like multiple copies running (like a keep alive restarter pattern). If ops has no access to the source to fix that or no one cares, then just run it in jail where you only get two processes, the restarter-forker and the forkee. Then hilarity can result if the restarter thinks the PID of the failed fork means something, like sending an email alert or logging the restart attempt. "Why are my logs now gigabytes of ERROR: restarted process new pid is -1?"
>>> resource.setrlimit(resource.RLIMIT_NPROC, (0, 0))
>>> os.fork()
Traceback (most recent call last):
File "<ipython-input-7-348c6e46312a>", line 1, in <module>
os.fork()
OSError: [Errno 11] Resource temporarily unavailableEvery system call can fail, even if it doesn't do something obvious like use disk resources. Ignoring this is how subtle bugs appear that seem unreproducible until you implement correct error handling.
It's not easy to guarantee that you deallocate on destructors exactly the resources that were allocated at the constructor when both of them can stop their execution at any time. Finally clauses are technically enough, but each allocation needs the same level of attention non-memory resources (e.g. connections, files) get on other languages.
> It's not easy to guarantee that you deallocate on destructors exactly the resources that were allocated at the constructor when both of them can stop their execution at any time.
I disagree; let's take this one case at time to keep it simple:
1. Destructors: within C++, if you're in a destructor, the object was fully constructed, and thus you know the exact set of resources requiring destruction. It is idiomatic C++ that a destructor should not throw; I'll discuss why below.
2. Constructors: these certainly can throw at any moment, as resource acquisition is often fraught with failures. That said, idiomatic C++ provides mechanisms (RAII, such as std::unique_ptr) to manage the partially constructed set of resources in a constructor, such that if something goes wrong, they will be automatically released by virtue of the variable going out of scope. Once you have the resource acquisition completed, you transfer ownership of the objects to the object you're constructing, which is practically guaranteed to be exception-free, since it's usually just moving a pointer under the hood.
> Finally clauses are technically enough
I don't really think you can both stand by the fact that destructors can throw at any moment and that finally clauses are enough, without making what amounts to an apples to oranges comparison. Take, for example, this function, where we assume releasing a resource can fail:
Foo() {
SomeResource resource;
// Assume the destruction of a SomeResource can fail.
// Other actions take place, some of which may raise/throw.
}
In this example, if the other actions throw an exception that causes Foo to itself abort, then SomeResource resource must be destructed. If we're assuming that destructor can also throw, we've now got two exceptions, and how do you handle two exceptions? (It's language dependent. Some discard an exception, some chain them, some, like C++, just terminate.)If we translate this to using some sort of "finally" construct, say in a garbage collected language:
def foo():
resource = aquire_some_resource()
try:
# other actions that may raise/throw.
finally:
resource.release() # but we're assuming this can also raise/throw.
You still have the same problem at the resource.release(): up to two exceptions can occur at a given point in the program, and you then need to know what your language does in that situation.The general gist of this is that if the "release" of some generic resource can fail, then you have to make harder decisions about what happens during a stack unwind due to some other error because now you have two errors. Do you ignore it? Log it? (can you log it?)
If releasing a resource cannot fail, destructors (and finally clauses in languages lacking RAII-style resource management) cannot fail.
It's still not particularly pleasant, but it is at least survivable and no information is lost.
"Should not", "practically". Your confidence is overwhelming. :-)
Exception safety in C++ may not be quite as much of a black art as it once was (say, before std::unique_ptr), but it is still something the programmer has to do, actively.
"If releasing a resource cannot fail, destructors (and finally clauses in languages lacking RAII-style resource management) cannot fail."
Yupper.
I'm probably unqualified to have an opinion on this, but I believe that the entire hatred for checked exceptions in Java comes from that general piece of idiocy and specifically from JDBC's urge to possibly throw a SQLException from close(). (Just what the hell is anyone supposed to do with that?)
pid,err = fork()
And what if $programmer forgets to check what's in err? What would pid contain in that case?
I mention this because I guess you quoted a kind of syntax that matches the one from Go.
So then I'm guessing that Go would simply ignore the error in this case.
However, having a proper exception mechanism, if you don't catch the problem, then it bubbles up, and the program doesn't continue with wrong data (which is a good thing => fail fast!).
This is the biggest downside of GoLang IMNSHO.
pid,_ = fork()
On second thought, it'd probably avoid some nasty production failures if that behaviour was true for everything which can return an error.
match fork() -> | Error(errno) -> ... | Pid(pid) -> ...
Or a general "Choice" sum, perhaps using phantom types so int<err> isn't compatible with int<pid>. But then all of a sudden, instead of a single word being returned, a tag and possibly variably-sized result has to be returned, and that's quite a hassle which doesn't fit well with C.
data ForkResult = Failure | Parent Int | ChildErlang gets away with it because dynamically typed + pattern matching + uses MRV for tagged returns, the second property makes it work.
If it's possible for the call to fail to return an object, you declare its type as optional (add a ? to it) and then the calling code has to explicitly "unwrap" the return value to get to the actual object - they can't simply go
var pid=fork()
pid.kill()
as that would be a compiler error. Instead, they need to go pid!.kill()
The idea is that they should check for the pid object being nil before doing the unwrapping. Of course, it's still possible for the coder to ignore that (just as it's possible, and depressingly common, for coders to catch and then ignore exceptions), but that's going to be a conscious decision because the compiler is telling them that there's a possible error condition here.If you haven't limited the number of processes a given non-root user can start to some value the machine can handle, sending SIGKILL to all of the user's processes is probably not going to do anymore damage.
If a program running as root doesn't correctly handle fork() failing, someone needs to be taken out back and beaten with a stick. Maybe the person who wrote the program, maybe the person who ran it as root. But somebody.
"killall -9 httpd" gave an unhelpful error message. "killall httpd" also gave an unhelpful error message. "killall", which would give you usage instructions in Linux, killed all processes on the system. Reading this article makes me figure that killall was likely a frontend to kill(-1, ...).
That day I learned a valuable lesson about reading man pages and understanding that not all unixes are the same.
I just recently finished a multithreaded program where I found obtaining the pid [on linux: getpid()] of child processes spawn was only effective by utilizing a common pipe that was non-blocking [fcntl(pipefd[1], F_SETFL, O_NONBLOCK) ].
In other, more humorous words, as a "parent," it's great to know what your "child," is doing (or in this sense), who your child is (the actual pid), instead of just kill SIGTERM them.
The simplest way I can think of is:
struct maybe {
bool isEmpty;
void* value;
}
Although I wonder if using C++ templates, classes and operator overloading is possible to make a more practical implementation (using void* does seem like a bad idea).In C ... hrm you could probably wrangle some macros around that struct if you were desperate.
if len(some_list) <= 0:
# Test for empty list
But it's just my way of covering my ass in case the laws of physics change during execution, or just in case weird bugs exist like those found in this article. * -1: error
* 0: success, in child
* > 0: success, in parent
Testing for <= 0 would cause you to think there's an error when you're really just the child process.Also the check as presumably written would never miss an error, it would just potentially assume valid return values were also errors.
Uh, no. If there is an error, there will be no child process but the parent process will think it is the child.
From the fork man page, emphasis mine: "On success, the PID of the child process is returned in the parent, and 0 is returned in the child."
"Also the check as presumably written would never miss an error, it would just potentially assume valid return values were also errors."
That was already discussed as a possibility; I was addressing the other. In the case you describe the software would never work at all, even when fork successfully forks, because the child will always think there was an error and presumably fall over rather than getting things done. That's probably the better case, in terms of development progress, because it would be spotted and fixed right away. But hopefully fixed correctly, and not converted to the broken-but-working-when-fork-succeeds other variant that also uses "<= 0".
There's no good substitute for reading the RETURN VALUE(S) section the manpage for every function and testing appropriately.
If you're not sure whether len() can return -1, and in turn what that value means, then you can't know that your code is any more correct.
In fact, this is going to be worse than a equals comparison because instead of having code that clearly doesn't handle a corner case you have code that lies about what values it can correctly handle. That makes it much harder to debug.
The article does not describe weird bugs - the behaviour it describes in fork() and kill() are by design, and well-documented. The real lesson here is to RTFM and understand what return values you get under what circumstances.
Sure, it's documented, but how often is it done on purpose? It seems like something that should at least be a separate function.
Hopefully doesn't happen often, but potentially very useful in narrow circumstances. The problem is it being -1, not it being available. If it took an argument to kill that you'd never accidentally generate as a PID then it might as well be a different function. Of course, with negative numbers otherwise referring to process groups, there's not a lot of room remaining, so yeah...
#include <unistd.h>
int main(void)
{
while(1) {
fork();
}
}But the forked processes will quit as soon as they are created. Thus, it won't work anyway.
"while(1) {fork()}; is "better", as the version you mention will quit as soon as it is not able to fork a new process."
or something similar
fork() or die "Cannot fork: $!";Thanks for providing an impetus to do so.
Using multiple returns for such situations is terrible, especially doing so in the way Go does.
(In this particular case that actually returns (for me) a bit of code that gets it right.)
reminds me of allocating memory for an error message to tell someone they are out of memory. :)
Imagine if C/POSIX had a checked ChildProcessCloneException