Reading /proc/pid/cmdline can hang forever
rachelbythebay.com
rachelbythebay.com
No, the problems start when you disable the kernel's OOM killer and try to act on your own.
Surprise, surprise: when you disable the time-tested built-in mechanisms that have been put to use in a multitude of cases that boggles the mind and try to implement your own, you're gonna have a bad time.
1) something like systemd starts a process in a cgroup, to limit memory usage
2) it disables the kernel OOM killer, and expects to
receive notification itself and kills the offending process
3) systemd crashes? This it can no longer process the
notifications, so you have a child process which is at
the memory limit but not killed.
This is a great example of why an init system should be as simple as possible and not die under any circumstance :) Why would your init system die? If you something like daemontools, they are designed to not even allocate memory after startup, because that could fail.Was she talking about systemd? Which process in systemd does the killing? Is it systemd-nspawn or something? (I'm not using systemd now.)
This has nothing to do with systemd.
And init of systemd doesn't die anyway, all the extra stuff is run in separate processes, process 1 is indeed, very simple.
It doesn't matter if it's not the PID 1 that fails. The article is saying: if the process that is supposed to receive OOM notification and kill the process fails, you will observe symptoms that are very hard to diagnose.
PID 1 dying is of course really bad, but if other processes in the init system die, your system can still be hosed, as pointed out here.
So? Not everything is about systemd.
systemd does not disable the OOM killer, so this has nothing to do with systemd.
You can create cgroups yourself you know - you don't need systemd for that.
It's time for an equivalent to Godwin's law for systemd.
The kernel isn't going to protect you from turning off critical bits.
Of course there's no reason for 'ps' not to build in its own timeout for i/o. It could cause premature failure on loaded boxes, but it wouldn't hurt anything.
1. oomkiller is on, and has no bugs in it. Great. We are all set.
2. oomkiller is off, process manager is on, and is bug free. Ditto.
3. oomkiller is on, but has bugs in it. Bad shit happens.
4. oomkiller is off, process manager is on, but has bugs. Bad shit happens.
Basically, the question is this: is it possible to develop a process manager that would do the job of oomkiller that has as high code quality as the Linux kernel. I am guessing the answer is yes, so we will always be hitting cases 1 and 2, or at least we'd be hitting cases 3 and 4 with roughly the same frequency.
Now, it's possible that a process manager can fail for reasons other than bugs. Not knowing enough about how cgroups, oom, etc., I can't say that you can write a 100% reliable (assuming it is 100% bug free) process manager without having kernel-level access. Perhaps there is something in the architecture of the whole thing that would prevent that.
I'm not trying to put the kernel devs on a pedestal or anything, but how can you compare a tool that's been tested in a myriad of use cases in literally everything from microcontrollers to clusters and big iron to something that's much less mature and probably won't see a tenth of the use cases? My bet is that you'll see more and more problems as people try to reinvent the wheel with their own "not-oom-killers", and blame it on the Linux kernel (ie, 99% of problems will be case 4, not nearly the same frequency). In that respect this article is incredibly informative as it will hopefully point people to check out their own code first. And the beauty is, if they do come up with a better oom-killer, they can always contribute a patch to the kernel.
Secondly, I doubt that a process manager with support for cgroups would need to run on microcontrollers. At least up to this point, I have not seen many microcontrollers running Linux containers.
Lastly, a hybrid solution could be good: process manager identifies what processes to kill and in what order, oomkiller does the killing.
I do agree with you about the hybrid solution, with a twist: since the oom-killer has been so widely tested, I would think that if someone needed different performance parameters, it might behoove them to start with the oom-killer and tune it, modifying the source if necessary, instead of re-inventing the wheel from scratch.
You can't assume the people who would replace oomkiller with their own process manager are the same ones who could write it well.
If something running within a cgroup can exceed what the cgroup allows then the cgroup is completely worthless as a concept.
Plus you have to assume some cgroups might be TRYING to DOS the machine its running on (e.g. shared hosting). If you give them a way to bypass the protections of cgroup and use up additional memory then they WILL take out the machine.
EDIT: I suspect that this problem is a consequence of a silly design that argv in the process holds the cmdline. Since apparently you can change the process name by changing argv, see:
http://www.uofr.net/~greg/processname.html
This seems to allow any process to mess with what the kernel sees as the cmdline. I hope there's not a "more" serious security issue hiding.
This is more of a suspicion than fact, I'm planning to look into the kernel code to get more info on that.
http://serverfault.com/questions/640248/ps-aux-hanging-on-hi...
I've yet to find the solution. Pretty much CentOS 6 install so no cgroups. It's pretty much IO resource starved and for the life of me can't figure out why it's only doing into /proc/pid# of a java process.
ISTM the bug is that things outside the memcg are waiting instead of getting ENOMEM.
Hackish test code to recurse infinitely and print the stack pointer until a segfault:
#include <stdio.h>
static inline unsigned long get_sp(void)
{
unsigned long ret;
__asm__("mov %%rsp, %0" : "=r" (ret));
return ret;
}
void main(void)
{
printf("%lu\n", get_sp());
main();
}
Running this and comparing the first and last values shows a stack size of about 8M.