Daemon Showdown: Upstart vs. Runit vs. Systemd vs. Circus vs. God
tech.cueup.com
tech.cueup.com
It's maddening that we have no clear separation of responsibilities for this stuff. And so we end up with things like Upstart's half-hearted logging, because it's not clear where Upstart's job stops.
I noticed this most of all when I was writing puppet manifests that ... create Upstart scripts.
Why do I have one tool that configures and supervises the bits at rest and a completely different tool that configures and supervises the same bits in flight?
Curious for people's thoughts on whether Solaris's SMF (or its general model) fits this bill.
I don't like Upstart, I don't agree that it's "admirably simple to use", it needlessly complex. For what I do, I don't need or want to care about runlevels, I always want respawning and I don't like the syntax of the configuration.
systemd seems like a really complex solution to a very simple problem, but I never used it.
Supervisord while being well documented and popular again seems to complicate what I want to do. Again I dislike the configuration, I fail to see why a shell script isn't better.
Runit I never used, because why would I replace daemon-tools with, what was when I first read about it, just a daemon-tools clone, made by people who didn't like Bernsteins licensing. I guess it's what I would most likely go for, if daemon-tools where not available.
The rest I never heard of.
I do take issue with your characterization of Upstart and Systemd as needlessly complex; almost all of that complexity is for their role as init.d replacements, and doesn't need to affect you if you just want to get a few programs running. Their big selling point, for me, is that one of them is usually the default and the config files are quite easy to write if you're not doing anything elaborate.
There are a couple nice runit features that stock daemontools lacks.
The ability to run each service in a separate process session with "runsvdir -P" is a big one. If your ./run file consists of a pipeline, bouncing the service with TERM will produce orphans otherwise. (I have some patches to daemontools-encore that do this on a per-service basis.)
The svlogd program from runit also looks nicer than multilog. I've tried to like tai64, I really have, but I'd rather just have human AND machine readable timestamps from the get go.
I like that service dependencies aren't statically declared in runit, but rather something you can block on in your ./run file.
http://smarden.org/pape/djb/manpages/
The manpages are installed by default in every system I use -- debian, ubuntu, MacPorts, and FreeBSD ports.
I agree with liking the simplicity of daemontools. I haved used it for a while, and it has served me well. I like that logging is split. I often run the log script as just a logger instance to get stdout to syslog.
Upstart is.... Not my favorite.
runit is nice. Like daemontools with manages and a few extra flags/features.
His recommendation is to use runit (based on djb's daemontools).
He recommends against using Bluepill, God, Foreman, supervisor and others.
[0] http://jtimberman.housepub.org/blog/2012/12/29/process-super...
However, I also think you'll be happy with Upstart or Systemd, even if you find them over-engineered and inelegant. They'll do what they need to with minimal configuration, and you're probably already using one of them behind the scenes. Why not use one of them, if you have it sitting right there, already installed and configured?
Nitpick: it's fine if your daemon forks, it just shouldn't background itself, which is different.
A nice succinct implementation of daemonization lives in the BSD sources: https://github.com/DragonFlyBSD/DragonFlyBSD/blob/master/lib...
But don't do this in your code ;) Do use a service manager to do this for you. Programs which lack a foreground option are harder to test and interact with due to the action-at-a-distance -- one of my chief objections with init.d-style service management over what I'll call daemontools-style service management. Nearly all popular daemons will have an option to stay in foreground for this reason. The only exceptional case that comes to mind is nginx, which does have a foreground option, but you'll lose zero-downtime upgrades if you use it (due to some extreme cleverness).
However, the flexibility of monit beyond process monitoring is really great. Anything from filesystem usage, cpu, memory to monitoring tcp ports, file checksums, you name it. I'm yet to find something you can't do with this little toy, and it's rock solid. I even use it to pull graphite stats[1] and report when certain thresholds are reached, like the percentage of 500s compared to other response codes in my nginx logs.
I also personally like its configuration syntax better than most yaml-like DSLs. It feels almost like writing natural text.
[1]http://blog.gingerlime.com/2013/graphite-alerts-with-monit/
If you use monit without runit, you end up in this weird world where monit starts and stops things in a very odd environment and it's flakier, environment variables are missing, etc. Also monit takes a long time to notice things are down, and doesn't start things right away on boot.
Finally if you don't boot your services with monit, you booted them with something else meaning that when a restart occurs, you're starting a service in a different environment than you booted in. weird.
so, let monit check health of things, let runit start and stop things.
which server (mongrel, thin, etc) do you recommend for playing nice with runit - which the OP mentioned has a problem with fork.
We're currently using supervisord but we've had to hack it to support more than 200 workers, and I'd like something with the ability to change the number of workers dynamically.
But the maintainers seem to have sub-zero interest in passing said listening sockets to multiple applications, which I'm sorry to see not be available. I'm thankful for this article, because it discusses Supervisord and Circus's willingness to do some pooled program handling.
It's a pain to learn, easy to screw up, non-trivial to test exhaustively, but provides highly flexible daemon migration and monitoring functionality you aren't likely to find anywhere else, for free, and for any imaginable service. (Actually you can monitor anything with it, including hardware. Check it out.)
Is pacemaker/corosync not more of a replacement for things like keepalived/heartbeatd (often used in conjunction with stonith, drbd), and as a way to run clusters of services?
You still need to launch and run the services themselves with something. (sysV init scripts, etc)
Even a cursory review of the docs seems to imply the same.
http://clusterlabs.org/doc/en-US/Pacemaker/1.1-crmsh/html/Pa...
http://clusterlabs.org/doc/en-US/Pacemaker/1.1-crmsh/html/Pa...
http://clusterlabs.org/doc/en-US/Pacemaker/1.1-crmsh/html/Pa...
I do see a reference to:
> Version 1 of Heartbeat came with its own style of resource
> agents and it is highly likely that many people have
> written their own agents based on its conventions.
> Although deprecated with the release of Heartbeat v2,
> they were supported by Pacemaker up until the release
> of 1.1.8 to enable administrators to continue to use
> these agents.
...so maybe that is what you were talking about.Resource agent scripts can support master/slave style services (including promotion/demotion) in addition to enabling the cluster to self-manage nontrivial overall system state transitions on a multi-host basis. You define the target running state with a declarative syntax that is replicated to across all nodes.
Heartbeat is an earlier platform that has now been functionally replaced by corosync.
Do you know of some decent article, blog post, video, whatever that presents these products, what they are good for and a typical use case?
Today I would use systemd with monit for additional monitoring. systemd ensures that the same environment is used regardless of how it is invoked, and the units are extremely simple. Monit can ensure a running server is behaving (memory limits, CPU limits, testing HTTP) and rely on systemctl for restarting processes.
Suggestions?
echo 'pidof myapp > /dev/null || su appuser -c "myapp --arguments" &' > ~/check_services.sh
echo 'exit 0' >> ~/check_services.sh
chmod 700 ~/check_services.sh
#Add ~/check_services.sh & to /etc/rc.local for startup
#Add */1 * * * /root/check_services.sh to crontabIt took us a long time to trace it down and we quickly switched to runit and supervisor on a few servers. All of our problems went away.
I am interested in any form of feedback on the issues you had. Falsely detecting a work crash sounds very weird and unprobable because Circus uses the system PID list to check on processes - so I wonder what happens in your case.
If you have a clean and useful taxonomy for this stuff, I'd be interested in hearing it.
2. Hard(-ish) resource limits and accounting: LXC w/ cgroups. Almost as good as full paravirtualization (Xen). There are still some issues with limiting resource contention impact between cgroups.
3. Softer resource limits, rogue app restarter: We've heavily modified bluepill because it seemed to lack insight on the needs and challenges of large-scale production ops. Specifically, we've added optional total child process limits (a few issues reported, fixed and even submitted a pull request). It might be useful to add max # of processes, nic bandwidth, iops and couple other checks.
Also worth considering:
4. Status monitoring: Icinga
5. Performance: collectd
6. Entropy injection: Chaos Monkey
It was designed to be simple (creating a new daemon conf file is trivial), low over-head (it's very init-ish in function so it needs to not take a lot of memory or processor time), stable (can't crash or it loses track of everything that it launched), and secure (lets you launch your daemons as different users--usually at a lower security level).
I've been using it myself an almost every server I administrate (I wrote it to scratch my own itch) for a couple years now and it's been very stable.
Zed Shaw has also starting a project of his own to tackle these things. Not sure how it handles multiple unrelated things. Seems to be a good idea to document in understandable Python code, what the requirements are to daemonize a process corrently.
That's weird, because that's one of the things i like about Upstart. The ability to set pam limits and others in the same file. Upstart seems to support the same options Runit memory wise: stack, data, memlock ...
limit nofile 10000 10000
limit nproc 1024 1024
nice 3
chroot /var/roots/mychroot
Has nothing to do with BSD style init? No part of the system should be either long or bash.
>it's own daemonisation in weird and wonderful ways is redundant and inconsistent
Calling daemon() is not that hard.
>and makes making something a daemon frustrating
That doesn't even make sense. People writing shitty software that should be a daemon but isn't makes that "frustrating".
supervisorctl update to start the new deamon without having to restart the whole supervisor process over again ...
I believe God and Monit have similar capabilities, although I haven't used them personally.
Or you can do something event more radical, like using fluentd[1], which looks really useful and well-designed. I like having this decision decoupled from the process manager. So, I would not count that as a strike against Angel.
Trololol.