Of course, this works better when you have many small processes rather than few monolithic ones. But now you're designing an Erlang system :)
---
Of course, this works better when you have many small processes rather than few monolithic ones. But now you're designing an Erlang system :)
---
I'm curious if this works in practice for you. The current OOM algorithm in Linux sums up the memory usage of a process and all its children. So there are good chances that the restarter process is killed first, and then the main software is killed too (when OOM killer realizes the last kill didn't free enough memory).
This is exactly the problem we're facing at work here: on a computational cluster, users sometimes start wild code that consumes all the memory, but the OOM decides to kill the batch-queue daemon first, because it's the root of all misbehaving processes. We have to explcitly set `oom_adj` on the important daemons to prevent the machines from becoming unresponsive because of a bad OOM decision.
Still, in your use-case, I'd definitely recommend only letting users run their "wild code" inside a memory cgroup+process namespace (e.g. an LXC container.)
Crash-only systems only work when a faulty component crashes itself before it crashes you. Processes modellable as mutually-untrustworthy agents should always have a failure boundary drawn between them. (User A shouldn't be able to bring down the cluster-agent; but they shouldn't be able to snipe user B's job by OOMing their job on the same cluster node, either.) And on a Unix box, the only true failure boundaries are jails/zones/containers; nothing else really stops a user from using up any number of not-oft-considered resources (file descriptors, PIDs, etc.)
I think it's surprisingly easy to get yourself in the situation where this is a concern for you[0] but you don't know how to solve it.
[0] Just run "adduser" and have SSH running, or just create an upstart job, or write a custom daemon that accepts and executes jobs from not-quite-trustworthy-undergrads, or...
It used to sum until Linux 2.6.36.
Diff'rent strokes. I also like the OOM killer, its a dastardly wonderful thing to tie lots of safety-critical things to .. and in the SIL-4 OS business (my domain), it is indeed a crash imperative to understand how to use the OOM properly. Or: not.
So in this light .. I know Erlang is "the thing" right now, but I feel I must just mention that:
>Of course, this works better when you have many small processes rather than few monolithic ones. But now you're designing an Erlang system :)
.. one could also be designing a Lua-based distribution, or JVM, or whatever you like, essentially, and integrating with oom_killer. There's nothing Erlang'y about it. Because if you're playing with the oom_killer, you're really making a distribution choice, in the topology.
If your app cares about oom, well kiddo .. you better not be doing anything less than excercising complete control over your distro, its launch policies, its use of the TextSegment as installed, and so on. Absolutely you're making Distribution decisions about the functionality of the combined system. oom_killer isn't useful by itself.
My point being, worrying about oom_killer isn't just something Erlang users need think about, nor are they the only ones who really 'get' why an oom_killer can be used nicely .. if you're building a distro, either for use as an embedded machine, a tight secure server image, or indeed even as a desktop user, well .. careful memory integration is, as you say, a harsh pass.
Incidentally, I use the oom_killer exactly as you mention, in a few embedded distro applications, specifically indeed, a kill of whatever 'lua' is hogging resources. Its an extraordinarily functional mechanism for recovery ..
I don't think Erlang is "the thing" right now. The "hot" languages for concurrent programming are Go and Node.js, for better or for worse.
ulimit is a better solution i.e. set reasonable constraints based on available resources.
Ok, it doesn't exactly support the quoted feature ("stop restarting if it exceeds X attempts in Y seconds"), but it does sleep for a second between restart attempts to mitigate the same problem: http://cr.yp.to/daemontools/supervise.html
Is your objection that daemontools is largely unused / unmaintained and lacks features (such as flapping-avoidance)? Or something else?
Just curious, thanks.
Either one of upstart, systemd, supervisord will give you equivalent solution that has more practical features. Daemontools was good enough a couple of years ago. (I used it then happily)
In short - there's nothing wrong with it really, but many alternatives are much better.
#! /bin/sh
sh $0 &
while true
do
perl -e 'push @big, 1 while 1' &
sh $0
doneThis isn't disastrous of course (replication + snapshots + redis aof), but still annoying.
Thus, the database itself will never degrade due to spilling to disk, but all other processes might. But if it's a DB box, that won't precisely matter.
They'll never swap out and bonus less memory used. (note transparent huge pages have only brought me pain so far)