monit: http://mmonit.com/monit
supervisord: http://supervisord.org
daemonize: http://bmc.github.com/daemonize
runit: http://smarden.org/runit/
perp: http://b0llix.net/perp/
DJB's daemontools: http://cr.yp.to/daemontools.html
systemd: http://www.freedesktop.org/wiki/Software/systemd
god: http://godrb.com
upstart: http://upstart.ubuntu.com
[1] http://0pointer.de/blog/projects/systemd-for-admins-4.html
[2] http://wiki.nginx.org/CommandLine#Upgrading_To_a_New_Binary_...
Anyone whose email provider provides persona automatically, or who runs their own identity provider, or who has taken 10 seconds to create an account on Mozilla's.
Besides, its control interface is the same as any other subsystem if you've configured it correctly -- i.e., "service nginx <verb>".
I agree it's not worth straining to make nginx's binary upgrade work with arbitrary supervision. However, if someone created a supervisor that solves this problem (systemd), I might give it a try.
It doesn't even have to be nginx's fault. It could be that some other process started fork-bombing the system and OOM-killer decided that killing nginx is the way to resolve it, before trying the actual offender.
oops, edit: pid 0 -> pid 1
sysv init!
All of my systems' processes are managed by it and have been for at least two decades.
Occasionally I do these periodic tasks as well, which are handled by a thing called "cron".
Yes this is sarcasm. There is a lot of wheel-reinventing done these days which is entirely unnecessary if you consider the long-forgotten "Unix philosophy".
myapp:234:respawn:/path/to/myscript
to your inittab, and sysvinit will relaunch it (if it dies).Regards.
I make sure I write processes that don't fail. This is done with tooling (valgrind, cachegrind, unit tests, integration tests) and testing (automated testing, experienced testers).
The whole process restart mechanism is a combination of string and sticky tape to abstract away the problem of half-arsed poorly written bits of functionality that crap themselves every few minutes.
I can and have written processes that stay alive for years (5 years 122 days pegged at 100% CPU being the record after which a reboot was required as the Sun Ultra 2 hosting it was replaced).
Please name one major, non-trivial network system software that has gone for 5 years, 122 days without a fault in production anywhere.
Engineers in other disciplines make best efforts not to fuck up. And then they accept that they, or others, will fuck up anyhow and they plan for that.
So do we. And so we should. Betting it all on Red 19 is a strategy of roulette, not serious work.
My point is that betting entirely on one strategy for mitigating faults is unnecessarily risky. Especially when additional levels of mitigation are easily installed and configured. In your other post you even point out a series of things that you do.
I don't see why process management is of a different kind.
Edit: removed unnecessary grandstanding.
I have written massively distributed systems which have an insanely high reliability requirement and it is really not the answer. I've been doing this for 25 years.
A more appropriate statement is:
Fail early and gracefully, recover always, expect failure.
Fail early - assertions up front. Prevention is better than cure.Fail gracefully - don't allow assertions to take the entire process out. Don't allow the language to crash the process out.
Recover always - design your system for recovery and understand recovery conditions.
Expect failure - know where and when something is going to fail and handle it.
Fail fast, fail often results in nasty shit like processes hung in restart cycles etc.
EDIT: I liked it so much (and it was so easy) that I wrote a blog post expounding how much I liked it and how to use it. http://moduscreate.com/monit-easy-monitoring/
I'd go on to mention various z/OS subsystems, but that's a bit esoteric even for HN :D
(Process management ties into my larger rant that nobody has properly combined it with configuration management. But nobody has time for this nonsense.)
The only thing you need to monitor is: whether a server answers the network request it was designed to. Outside of that you might optionally want to know whether the disk is full, ram is maxed (thus putting linux into swap) or if the cpu runs too high to cope with losing some servers at peak, but really that's all optional if you're in ec2 and can just spin up more servers on a moment's notice.
You can gather all this data for yourself with Newrelic or if you want you can send data to graphite or if you're old-fashioned you can use Icinga in place of Nagios because it keeps history in a database. If the developers want to know about the process for the application they implemented you can put Newrelic on the server for them, and put the system Newrelic thing on there too, just don't pay attention to it or pretend it's important until something breaks.
The important catch here, the thing that is critical to this whole line of thinking: you have to have thought things through before you built them, focused on having one service per os and real redundancy throughout the environment, and then critically your kick should be fast enough that if a server has some kind of problem in production you don't fix it you just re-kick it. That means your kick throws the os on there, then triggers salt or ansible or chef to configure every single detail and then triggers a deploy of internally-developed applications. That means you have to test the kick to death before you can rely on it to rebuild you something live. If the problem is recurring you can use immediate tools, jdump or whatever, to get some data, give it to the application's developers, and let them try to recreate it in staging while you go ahead and re-kick the prod server and go back to writing documentation for lesser ops to not read, drinking at your desk, reading hackernews, acting as a cia listening post for cat pictures or whatever else passes the time.
1. daemontools and runit are practically identical. I do prefer runit somewhat, as svlogd has a few more features than multilog (syslog forwarding, more timestamp options), and sv has a few more options than svc (it can issue more signals to the supervised process).
2. Among the criteria I look for in a process manager are: (1) the ability to issue any signal to a process (not just the limited set of TERM and HUP), and (2) the ability to run some kind of test as a predicate for restarting or reloading a service. The latter is especially useful to help avoid automating yourself into an outage. As far as I'm aware, none of the above process supervisors can do that, so I tend to eschew them in favor of initscripts and prefer server implementations that are reliable enough not to need supervision.
[1] http://www.freedesktop.org/software/systemd/man/systemd.serv...
There was a nasty race condition for a while that locked circus up, and it wouldn't restart crashed procs. For a while, you couldn't specify a timeout on the commandline - on a major version, too. It was 1.0.0 in master for a while, and then went backwards to 0.7.0. We were sitting on master for the fix to the aforementioned timeout, and so no updates happened until we realized what happened and then manually "downgraded."
All in all, it really feels like we're either using it wrong (probably, we're adding & removing processes on the fly), or we're the only ones really loading it up with a ton of processes which may or may not flap a lot.
If you don't mind, why are you moving away from supervisord?
That said, the ability to manage sockets sounds very interesting, hopefully simplifying my stack even more (my current use case is in getting the most performance out of a small VPS, so removing things from the stack would hopefully clear up RAM for actual web workers). I've been running a small site using nginx->circus->chaussette->django and it's faster (and was simpler to configure) than my standard nginx->supervisord->uwsgi->django deployment.
Also, supervisord isn't without its issues, and one of them is managing processes at scale (see https://github.com/Supervisor/supervisor/issues/26, that bug is 2 years in the making). Circus at least claims they intend to support thousands (via http://circus.readthedocs.org/en/0.9.2/rationale/) so I'd be interested in seeing what they bring to the table on that front.
This was fixed in 0.7.1
> For a while, you couldn't specify a timeout on the commandline - on a major version, too.
To my knowledge this was never released.
> It was 1.0.0 in master for a while, and then went backwards to 0.7.0.
Yes we decided for a while the next version would be 1.0 then we changed our mind. All happened in master and was not released, so I don't see the problem here;
> All in all, it really feels like we're either using it wrong (probably, we're adding & removing processes on the fly), or we're the only ones really loading it up with a ton of processes which may or may not flap a lot.
I am still available for any help. Circus is young but works for our needs. If you are happily using Supervisord, that's fine - but keeping on posting your negative experience on HN from 3 months ago without having tried the tool recently --while we addressed to my knowledge all the bugs you mention-- is a bit inapropriate imho
> This was fixed in 0.7.1
We're still using Circus under production loads, and we're still seeing it go unresponsive and chew through a ton of CPU. Unfortunately we haven't been able to reliably reproduce it, so until we can, it's not something we can fix.
> To my knowledge this was never released.
It's happened a couple times, the second one was likely just on master, but the first was happening from a pip install. https://github.com/mozilla-services/circus/issues/457 https://github.com/mozilla-services/circus/issues/380
> If you are happily using Supervisord, that's fine
We're not, we're still using Circus while we find or build something else.
Can systemd do that? I like everything else I've heard about it, but I haven't seen any documentation regarding this sort of usage.
For most standalone systems though, Upstart works nicely enough as long as services don't need too much coordination.
systemd is currently adding support for dynamic service "instances", which would allow you to launch and close such services on the fly.
So, depending on how dynamic you need the list of virtual services to be, the answer is either "yes, that works right now" or "yes, that's available in the latest systemd release, but probably not in your distro yet".
My main concern is that systemd handles too much and I feel it's going to be hard to change. I guess we can start using systemd without removing supervisord and move from there.
The code is designed to be shot, and will always recover after a restart. rundir/svc will neatly reap a process and re-start it again. And can be used separately.
Also I'll put in a shameless plug for my side project, a service management tool for multi project development; Hack On[2]
For reference see this bug which has been open for 2 years.
I think it's a bit unfair to downvote me, the OP has posted a Poll for what people use, not what they recommend. We use God, so I've asked for God to be added to the list.
I've been dogfooding it in a production environment for a couple years and it's been pretty solid.
I have also used god at some point, but I kept having trouble. I can't remember exactly what was wrong but it never quite worked correctly for me. Probably PEBCAK.
http://quickhist.onloop.net/monit=75,supervisord=107,daemoni...
I haven't looked recently at alternatives, but I'm open to it.
Unless you did really mean process and not daemon, then it's supervisord.
We use this to monitor services at the application level.
Systemd was pretty stable until user mode flat out broke in 205. I use it to manage my entire desktop session.
Supervisor programs aren't just about restarting processes if they die. They manage dependencies, reloading, a central authority for starting/stopping and other administrative tasks. Some additionally do things like logging.
And besides, even the most stable software can't be guaranteed bug-free (though I do agree that a better action might be to stop and yell loudly rather than blindly restart). Why forego a safety net?
People have stopped writing software designed to run continuously.
Some notes from 20 years of Unix fudging:If a process dies, it's a problem with the process, not an external problem. Make it reliable. Software shouldn't fail at all. PostgreSQL, Postfix, init never crash from experience. Why shouldn't your software do the same?
Reloading: system-wide: "service whatever restart/reload"
Dependencies: Dependencies are hell. Having used BIG Unix kit with power sequencers and stuff, I wouldn't ever take on a dependency. If I did, the dependent processes should fail gracefully and retry.
Central authority: service xxx restart/reload again...
Logging: syslog? On our kit, we use Splunk but it's pretty much overkill for most people.
If you can't sleep at night because your processes crash, there's no good employing someone (a supervisor) to go around and do CPR on them. Fix the root problem.
We pick software that is reliable and trustworthy then test it thoroughly.
Those are forgotten arts outside of the enterprise. Elsewhere, any new technology that falls from the sky is picked whether it works properly or not.
We mitigate hardware issues with either hot spares or clustering. We have three datacentres distributed geographically with entire redundant sets of equipment (4x 42U racks in each).
With respect to software failures, we test everything thoroughly including failover conditions etc. Everything is load tested as well.
We are prepared for emergency. We have dedicated people ready to jump on that.
However preventing these things ever being needed is a professional responsibility which is my point.